<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://sokwe.janegoodall.org/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=StevanEarl</id>
	<title>sokwedb - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://sokwe.janegoodall.org/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=StevanEarl"/>
	<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/wiki/Special:Contributions/StevanEarl"/>
	<updated>2026-09-30T21:39:55Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.35.6</generator>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=825</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=825"/>
		<updated>2026-09-30T19:33:35Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Problem #266. Documents a code defect (not a conversion problem); added here since the code is static following the handoff to Ode&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commits efa75af037a5e22b2b43675b435f472193a6672f and 921b5ca3e9ca3d1e06fe189b4775aae7dcc484bf&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
9/3/2026 - ICG: Actually this returns cases where a scientific food name is associated with more than one local food name. That is, There can be multiple words for the same latin name.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt; will be derived from&lt;br /&gt;
&amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt;, not &amp;lt;code&amp;gt;fl_sci_food_name&amp;lt;/code&amp;gt;.  Problem&lt;br /&gt;
#87 therefore controls description collisions for this conversion.&lt;br /&gt;
&lt;br /&gt;
The historical query above does not test the relationship stated in the&lt;br /&gt;
heading: it finds a scientific name shared by multiple local names.  Do not use&lt;br /&gt;
the historical count of 47 as a conversion assertion.  If this issue is&lt;br /&gt;
revisited, first replace the query with one that tests the intended&lt;br /&gt;
relationship against the refreshed source.&lt;br /&gt;
&lt;br /&gt;
== (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
9/4/2026 &lt;br /&gt;
Ian fixed mifupa, miti and mizizi in FOOD_BOUT by changing to UNRECORDED. DELETED FROM FOOD_PART_LOOKUP&lt;br /&gt;
Consolidated insects to &amp;#039;dudu&amp;#039;&lt;br /&gt;
changed all &amp;quot;NA&amp;quot; to &amp;quot;None&amp;quot;&lt;br /&gt;
Kept unrecorded&lt;br /&gt;
fixed spellings of utomvi and chipukizi&lt;br /&gt;
&lt;br /&gt;
I made all associated changes in FOOD_BOUT, choosing to use names rather than initials&lt;br /&gt;
&lt;br /&gt;
== (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Use &amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt; for&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt;.  Exclude lookup rows where&lt;br /&gt;
&amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt; is true before deriving descriptions.  Leave each&lt;br /&gt;
noncolliding generalized description unchanged.  When one generalized&lt;br /&gt;
description is shared by multiple eligible local names, derive each target&lt;br /&gt;
description as:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
fl_sci_food_name_gen || &amp;#039; -- &amp;#039; || fl_local_food_name&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This representation is stable and preserves both source values.  Do not use&lt;br /&gt;
mung&amp;#039;s row-order-dependent numeric suffixes.  In the 2026-09-13 snapshot, 39&lt;br /&gt;
generalized descriptions remained shared by 111 eligible local names after&lt;br /&gt;
applying &amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt;.  Rerun a corrected collision query as this&lt;br /&gt;
problem is encountered; the historical count of 42 predates the approved&lt;br /&gt;
exclusion rule.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Do not silently choose between a compound primary part and a conflicting&lt;br /&gt;
explicit second part.  If such a conflict remains after refreshed source&lt;br /&gt;
cleanup, exclude the exact source row and document all source columns in the&lt;br /&gt;
loader predicate.&lt;br /&gt;
&lt;br /&gt;
The dated 2026-09-13 source contained one conflict:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Date !! Focal !! Begin !! End !! Primary part !! Primary name !! Explicit second part !! Second name&lt;br /&gt;
|-&lt;br /&gt;
| 1985-10-26 || EV || 11:22 || 11:51 || MATUNDA; CHIPUKIZI || MBULA || MATUNDA; CHIPUKIZI || BISHURUSHURU&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Under the approved parser, the compound second token is&lt;br /&gt;
&amp;lt;code&amp;gt;CHIPUKIZI&amp;lt;/code&amp;gt;, while the explicit second value is the entire compound&lt;br /&gt;
&amp;lt;code&amp;gt;MATUNDA; CHIPUKIZI&amp;lt;/code&amp;gt;.  The older table above uses the stale spelling&lt;br /&gt;
&amp;lt;code&amp;gt;CHIPUKIZA&amp;lt;/code&amp;gt;.  Reproduce the complete current row exactly before&lt;br /&gt;
adding the exclusion; do not rely on this dated spelling or count.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Add &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt;, described as &amp;lt;code&amp;gt;No data recorded&amp;lt;/code&amp;gt;, to both&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;FOOD_PARTS&amp;lt;/code&amp;gt;.  Create a Seq 2 row when&lt;br /&gt;
either approved secondary component exists.  Use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; only for&lt;br /&gt;
the missing half of that pair:&lt;br /&gt;
&lt;br /&gt;
* part exists, name missing: use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; for FoodName;&lt;br /&gt;
* name exists, part missing: use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; for FoodPart;&lt;br /&gt;
* neither exists: create no Seq 2 row; and&lt;br /&gt;
* both exist: preserve both approved values.&lt;br /&gt;
&lt;br /&gt;
Normalize colon and semicolon delimiters and derive ordered compound tokens for&lt;br /&gt;
Seq 1 and Seq 2.  If a compound-derived second part conflicts with an explicit&lt;br /&gt;
second part, exactly exclude the row under Problem #89.  Do not silently apply&lt;br /&gt;
precedence.  The dated source had 12 secondary names without an explicit second&lt;br /&gt;
part; these are eligible for the approved &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; FoodPart rather&lt;br /&gt;
than exclusion.  Rerun both queries above with the approved clean normalization&lt;br /&gt;
and exclusion projection as this problem is encountered.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, add one deterministic spelling of each case-insensitive GROOM_SCANS extractor value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite gs_extracted_by to the exact spelling stored in clean.people.&lt;br /&gt;
&lt;br /&gt;
The ordinary support-table loader copies these rows to codes.people with active set to true before B-record groom scans are loaded. This preserves all 44,679 source rows and leaves the production foreign key and active-person rule intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of `U` in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-11.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#136) Some GROOM_SCANS direction codes have trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 GROOM_SCANS records where GS_direction is &amp;#039;G &amp;#039; instead of &amp;#039;G&amp;#039;. The trailing Access padding prevents an exact match with the valid one-character direction code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in the clean schema, so query the easy schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;#039;&amp;quot;&amp;#039; || gs_direction || &amp;#039;&amp;quot;&amp;#039; AS untrimmed_direction,&lt;br /&gt;
       &amp;#039;&amp;quot;&amp;#039; || BTRIM(gs_direction) || &amp;#039;&amp;quot;&amp;#039; AS trimmed_direction,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM easy.groom_scans&lt;br /&gt;
 WHERE gs_direction IS DISTINCT FROM BTRIM(gs_direction)&lt;br /&gt;
 GROUP BY gs_direction,&lt;br /&gt;
          BTRIM(gs_direction)&lt;br /&gt;
 ORDER BY gs_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This reports 3 rows having &amp;quot;G &amp;quot;, which normalizes to &amp;quot;G&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Trim leading and trailing whitespace from GS_direction in the clean schema. This preserves all source rows and allows the direction codes to be mapped normally by the production loader.&lt;br /&gt;
&lt;br /&gt;
Resolved by commit 3ad4f5e1e2355895e5e465c4b8049197c8dce295&lt;br /&gt;
&lt;br /&gt;
== * (#137) GROOM_BOUT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A grooming event must relate to a WATCHES row. Normal B-record watches are created from clean.follow, but 1,536 current GROOM_BOUT rows have no follow with the same focal and date after animal-ID whitespace normalization. GROOM_BOUT has no community column, so the loader cannot create a watch directly from the source row.&lt;br /&gt;
&lt;br /&gt;
Of these rows, 1,399 have exactly one COMMUNITY_MEMBERSHIP row covering the observation date. The remaining 137 do not have exactly one dated membership and must not be assigned a community by inference.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid = gb.grm_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = gb.grm_fol_date)&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The temporarily excluded subset is identified by:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid = gb.grm_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = gb.grm_fol_date)&lt;br /&gt;
       AND 1 &amp;lt;&amp;gt; (&lt;br /&gt;
         SELECT count(*)&lt;br /&gt;
           FROM clean.community_membership&lt;br /&gt;
          WHERE community_membership.cm_b_animid =&lt;br /&gt;
                  gb.grm_fol_b_focal_animid&lt;br /&gt;
                AND gb.grm_fol_date BETWEEN&lt;br /&gt;
                      community_membership.cm_start_date&lt;br /&gt;
                      AND community_membership.cm_end_date)&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
As established for ad-hoc observations in Problem #128, prefer an existing B-record WATCHES row and otherwise reuse an existing Other watch for the focal and date. When no watch exists and exactly one COMMUNITY_MEMBERSHIP interval covers the observation date, create an Other watch using that membership&amp;#039;s community. This preserves the grooming bout without falsely representing it as part of a follow.&lt;br /&gt;
&lt;br /&gt;
Do not infer a community from unrelated observations when there is not exactly one dated membership. Temporarily exclude those 137 rows from the production groom-bout loader. Three overlap Problem #138, so this exclusion adds 134 rows to the excluded union. Investigators must determine the focal&amp;#039;s community on the observation date or correct the missing follow or membership data in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#138) GROOM_BOUT initiator or terminator is not a member of the grooming pair ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After applying the approved Problem #112 interpretation of U as unknown and trimming animal IDs under Problem #139, 2,149 GROOM_BOUT rows have a nonempty initiator or terminator animal ID that is neither the focal nor the recorded grooming partner. There are 1,776 initiator mismatches and 2,133 terminator mismatches, with overlap between those sets.&lt;br /&gt;
&lt;br /&gt;
Production GROOMINGS.Initiator and GROOMINGS.Terminator values must reference ROLES.PID rows belonging to participants in the same grooming event. The source does not establish whether the mismatched value, the recorded partner, or another field is incorrect, so the loader cannot safely choose a role or add another participant.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
   gb.grm_fol_b_focal_animid,&lt;br /&gt;
   gb.grm_time_begin,&lt;br /&gt;
   gb.grm_b_partner_animid,&lt;br /&gt;
   gb.grm_direction,&lt;br /&gt;
   invalid_reference.source_column,&lt;br /&gt;
   invalid_reference.animid AS offending_animid,&lt;br /&gt;
   gb.grm_extracted_by,&lt;br /&gt;
   gb.grm_problems,&lt;br /&gt;
   gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
   CROSS JOIN LATERAL (&lt;br /&gt;
     VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
        (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
   ) AS invalid_reference(source_column, animid)&lt;br /&gt;
 WHERE invalid_reference.animid IS NOT NULL&lt;br /&gt;
   AND invalid_reference.animid NOT IN (&lt;br /&gt;
     gb.grm_fol_b_focal_animid,&lt;br /&gt;
     gb.grm_b_partner_animid)&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
      gb.grm_fol_b_focal_animid,&lt;br /&gt;
      gb.grm_time_begin,&lt;br /&gt;
      gb.grm_b_partner_animid,&lt;br /&gt;
      invalid_reference.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows from the production groom-bout loader. The project investigators must determine the correct grooming partner and the correct initiator or terminator in Access. Keep the production requirement that an initiator or terminator reference a participant role in the same grooming event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#139) GROOM_BOUT animal IDs contain edge whitespace ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 50 animal-ID values with leading or trailing whitespace across 49 GROOM_BOUT rows: 36 focal values, 9 partner values, 2 initiator values, and 3 terminator values. The first runtime failure used partner value &amp;lt;code&amp;gt;BE &amp;lt;/code&amp;gt;, which does not reference BIOGRAPHY_DATA even though the trimmed value &amp;lt;code&amp;gt;BE&amp;lt;/code&amp;gt; does.&lt;br /&gt;
&lt;br /&gt;
All 36 focal and all 9 partner values identify existing BIOGRAPHY rows after trimming. Trimming is the established lossless conversion treatment for animal-ID edge whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying easy because the values are normalized in clean&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       animal_id.source_column,&lt;br /&gt;
       animal_id.animid,&lt;br /&gt;
       BTRIM(animal_id.animid) AS trimmed_animid&lt;br /&gt;
  FROM easy.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_FOL_B_focal_AnimId&amp;#039;, gb.grm_fol_b_focal_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_partner_AnimId&amp;#039;, gb.grm_b_partner_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS animal_id(source_column, animid)&lt;br /&gt;
 WHERE animal_id.animid IS DISTINCT FROM BTRIM(animal_id.animid)&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          animal_id.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Trim leading and trailing whitespace from the four GROOM_BOUT animal-ID columns in the clean schema. Apply the Problem #112 conversion of initiator or terminator U values to SQL NULL after trimming. Do not change GRM_direction.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#140) GROOM_BOUT focal or partner animal IDs are absent from BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After the lossless Problem #139 whitespace normalization, 412 GROOM_BOUT rows have a focal or grooming partner that is absent from BIOGRAPHY: 131 rows have an absent focal and 337 have an absent partner. Some rows occur in both counts. Production WATCHES.AnimID and ROLES.Participant values must reference BIOGRAPHY_DATA.AnimID.&lt;br /&gt;
&lt;br /&gt;
The first otherwise eligible runtime failure is the 1992-10-13 PF/MGE bout. The source does not identify a production animal for MGE, so no sentinel or other animal can be substituted safely.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       missing_participant.source_column,&lt;br /&gt;
       missing_participant.animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_FOL_B_focal_AnimId&amp;#039;, gb.grm_fol_b_focal_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_partner_AnimId&amp;#039;, gb.grm_b_partner_animid)&lt;br /&gt;
       ) AS missing_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography&lt;br /&gt;
          WHERE biography.b_animid = missing_participant.animid)&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          missing_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows from the production groom-bout loader. Of the 412 matches, 206 overlap Problems #137 or #138 and 206 are newly excluded. Investigators must identify the intended animals and correct the Access data. Keep the production foreign keys to BIOGRAPHY_DATA intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#141) GROOM_BOUT participants are observed before entry into the study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Seven GROOM_BOUT rows involve a focal or partner before that animal&amp;#039;s BIOGRAPHY.EntryDate: one focal occurrence and six partner occurrences. Production rejects a role dated before its participant entered the study. The first otherwise eligible failure observes SN on 1996-03-11, before its 1996-06-03 entry date.&lt;br /&gt;
&lt;br /&gt;
The identifying predicate uses a strict less-than comparison, so an observation on EntryDate remains valid.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       early_participant.source_column,&lt;br /&gt;
       early_participant.animid,&lt;br /&gt;
       biography.b_entrydate,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_FOL_B_focal_AnimId&amp;#039;, gb.grm_fol_b_focal_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_partner_AnimId&amp;#039;, gb.grm_b_partner_animid)&lt;br /&gt;
       ) AS early_participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography&lt;br /&gt;
         ON biography.b_animid = early_participant.animid&lt;br /&gt;
 WHERE gb.grm_fol_date &amp;lt; biography.b_entrydate&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          early_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows from the production groom-bout loader. One overlaps Problems #137, #138, or #140, so this exclusion adds six rows to the excluded union. Investigators must correct the observation date, participant, or biography entry date in Access. Keep the production rule and its inclusive EntryDate boundary intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#142) GROOM_BOUT participants are observed after departure from the study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Sixty GROOM_BOUT rows involve a focal or partner after that animal&amp;#039;s BIOGRAPHY.DepartDate: five rows have an affected focal and 56 have an affected partner. Some rows occur in both counts. Production rejects a role dated after its participant departed the study. The first otherwise eligible failure observes FI on 1997-07-30, one day after its 1997-07-29 departure date.&lt;br /&gt;
&lt;br /&gt;
The identifying predicate uses a strict greater-than comparison, so an observation on DepartDate remains valid.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       late_participant.source_column,&lt;br /&gt;
       late_participant.animid,&lt;br /&gt;
       biography.b_departdate,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_FOL_B_focal_AnimId&amp;#039;, gb.grm_fol_b_focal_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_partner_AnimId&amp;#039;, gb.grm_b_partner_animid)&lt;br /&gt;
       ) AS late_participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography&lt;br /&gt;
         ON biography.b_animid = late_participant.animid&lt;br /&gt;
 WHERE gb.grm_fol_date &amp;gt; biography.b_departdate&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          late_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows from the production groom-bout loader. Eleven overlap Problems #137, #138, #140, or #141, so this exclusion adds 49 rows to the excluded union. Investigators must correct the observation date, participant, or biography departure date in Access. Keep the production rule and its inclusive DepartDate boundary intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#143) GROOM_BOUT other-partner flags have non-boolean values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
GROOMINGS.Others is a required boolean, while GROOM_BOUT.GRM_other_partners contains several encodings. The established Access boolean-style convention treats NULL or the empty string as false, and the source also contains lowercase y and n values that differ only by case. There are 4,266 blank or NULL values and 561 lowercase values.&lt;br /&gt;
&lt;br /&gt;
After those lossless boundary normalizations, nine rows retain values that cannot be mapped to a boolean: B in three rows, G in one row, M in two rows, and U in three rows. The first otherwise eligible runtime failure has U on the 2002-07-28 SA/SR bout.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary before clean-schema normalization&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(grm_other_partners, &amp;#039;&amp;amp;lt;NULL&amp;amp;gt;&amp;#039;) AS raw_value,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM easy.groom_bout&lt;br /&gt;
 GROUP BY grm_other_partners&lt;br /&gt;
 ORDER BY raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;unmappable values in clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_other_partners,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
 WHERE gb.grm_other_partners NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert NULL and empty GRM_other_partners values to N and normalize lowercase y and n to uppercase. Temporarily exclude the nine rows whose normalized value is not Y or N. One overlaps Problems #137, #138, #140, #141, or #142, so this exclusion adds eight rows to the excluded union.&lt;br /&gt;
&lt;br /&gt;
Investigators must determine whether B, G, M, and U mean that other grooming partners were present and correct the Access data. Keep GROOMINGS.Others required and boolean.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#144) GROOM_BOUT times are outside the production observation window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Six GROOM_BOUT rows have start and stop times outside the production EVENTS range of 04:00 through 20:00 inclusive. All six violate both the start and stop constraints. The first runtime failure is the 2007-03-28 NUR/KS bout from 02:55 through 02:57.&lt;br /&gt;
&lt;br /&gt;
The source does not establish whether these are valid nighttime observations or mistyped times. Changing the production observation window requires a separate decision.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_time_end,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
 WHERE gb.grm_time_begin &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
       OR gb.grm_time_begin &amp;gt; &amp;#039;20:00&amp;#039;::TIME&lt;br /&gt;
       OR (gb.grm_time_end IS NOT NULL&lt;br /&gt;
           AND (gb.grm_time_end &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
                OR gb.grm_time_end &amp;gt; &amp;#039;20:00&amp;#039;::TIME))&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the six rows from the production groom-bout loader. None overlap Problems #137, #138, #140, #141, #142, or #143. Investigators must confirm the intended times or decide separately whether the production observation window should change. Keep the current production constraints intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#145) GROOM_BOUT records the same animal as focal and partner ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Four GROOM_BOUT rows identify the same animal as both the focal and grooming partner. Production represents grooming as a dyadic event and requires each ROLES participant to be unique within an event, so the second role violates the unique participant-and-event constraint. The first failing row is the 2011-07-31 TOM/TOM bout, whose source problems text includes &amp;quot;WRONG FOCAL&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
 WHERE gb.grm_fol_b_focal_animid = gb.grm_b_partner_animid&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the four rows from the production groom-bout loader. None overlap Problems #137, #138, #140, #141, #142, #143, or #144. Investigators must identify the intended focal or grooming partner and correct the Access data. Keep the production requirement that a participant occur at most once in an event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#146) GROOM_SCANS participants are observed after departure from the study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Sixty-four GROOM_SCANS rows involve a grooming participant after that animal&amp;#039;s BIOGRAPHY.DepartDate: eight rows have an affected &amp;lt;code&amp;gt;GS_B_chimp1_AnimId&amp;lt;/code&amp;gt; and 56 have an affected &amp;lt;code&amp;gt;GS_B_chimp2_AnimId&amp;lt;/code&amp;gt;. Production rejects a role dated after its participant departed the study. The first runtime failure observes MT on 1978-05-23, after its 1974-11-01 departure date.&lt;br /&gt;
&lt;br /&gt;
The identifying predicate uses a strict greater-than comparison, so an observation on DepartDate remains valid. Seventeen participant occurrences are on their departure date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       late_participant.source_column,&lt;br /&gt;
       late_participant.animid,&lt;br /&gt;
       biography.b_departdate,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GS_B_chimp1_AnimId&amp;#039;, gs.gs_b_chimp1_animid),&lt;br /&gt;
                (&amp;#039;GS_B_chimp2_AnimId&amp;#039;, gs.gs_b_chimp2_animid)&lt;br /&gt;
       ) AS late_participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography&lt;br /&gt;
         ON biography.b_animid = late_participant.animid&lt;br /&gt;
 WHERE gs.gs_date &amp;gt; biography.b_departdate&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid,&lt;br /&gt;
          late_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the 64 affected rows from the production B-record groom-scan loader. Investigators must correct the observation date, participant, or biography departure date in Access. Keep the production rule and its inclusive DepartDate boundary intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#147) GROOM_SCANS focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A B-record groom scan must relate to a B-type WATCHES row. There are 673 current GROOM_SCANS rows whose focal and date have no matching clean.follow row. Of these, 665 rows have exactly one COMMUNITY_MEMBERSHIP interval covering the observation date and eight rows, comprising three focal/date keys, have no dated membership.&lt;br /&gt;
&lt;br /&gt;
The first runtime failure is the 1978-06-29 ST scan at 14:15. Its source table identifies it as B-record data, and ST has exactly one community membership on that date, so the community is unambiguous even though the corresponding follow is absent.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid = gs.gs_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = gs.gs_date)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The temporarily excluded subset is identified by:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid = gs.gs_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = gs.gs_date)&lt;br /&gt;
       AND 1 &amp;lt;&amp;gt; (&lt;br /&gt;
         SELECT count(*)&lt;br /&gt;
           FROM clean.community_membership&lt;br /&gt;
          WHERE community_membership.cm_b_animid =&lt;br /&gt;
                  gs.gs_fol_b_focal_animid&lt;br /&gt;
                AND gs.gs_date BETWEEN&lt;br /&gt;
                      community_membership.cm_start_date&lt;br /&gt;
                      AND community_membership.cm_end_date)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Reuse an existing B-type WATCHES row when one was created by an earlier loader. Otherwise, when exactly one COMMUNITY_MEMBERSHIP interval covers the observation date, create a B watch using that membership&amp;#039;s community. The GROOM_SCANS source identifies these observations as B-record data, so this does not invent the watch type.&lt;br /&gt;
&lt;br /&gt;
Do not infer a community when no dated membership exists. Temporarily exclude the eight unresolved rows; they do not overlap Problem #146. Investigators must determine the focal&amp;#039;s community on the observation date or correct the missing follow or membership data in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#148) GROOM_SCANS participants are observed before entry into the study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Four GROOM_SCANS rows involve a grooming participant before that animal&amp;#039;s BIOGRAPHY.EntryDate. All four affected values are in &amp;lt;code&amp;gt;GS_B_chimp2_AnimId&amp;lt;/code&amp;gt;. Production rejects a role dated before its participant entered the study. The first runtime failure observes KP on 1981-09-15, before its 1997-03-16 entry date.&lt;br /&gt;
&lt;br /&gt;
The identifying predicate uses a strict less-than comparison, so an observation on EntryDate remains valid. No current GROOM_SCANS participant is observed exactly on its entry date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       early_participant.source_column,&lt;br /&gt;
       early_participant.animid,&lt;br /&gt;
       biography.b_entrydate,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GS_B_chimp1_AnimId&amp;#039;, gs.gs_b_chimp1_animid),&lt;br /&gt;
                (&amp;#039;GS_B_chimp2_AnimId&amp;#039;, gs.gs_b_chimp2_animid)&lt;br /&gt;
       ) AS early_participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography&lt;br /&gt;
         ON biography.b_animid = early_participant.animid&lt;br /&gt;
 WHERE gs.gs_date &amp;lt; biography.b_entrydate&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid,&lt;br /&gt;
          early_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the four affected rows from the production B-record groom-scan loader. They do not overlap Problems #146 or #147. Investigators must correct the observation date, participant, or biography entry date in Access. Keep the production rule and its inclusive EntryDate boundary intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#149) GROOM_SCANS times are outside the production observation window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Sixty-one GROOM_SCANS rows have times outside the production EVENTS range of 04:00 through 20:00 inclusive. The values range from 00:00 through 23:50. The first runtime failure is the 1982-01-13 PI scan at 03:30.&lt;br /&gt;
&lt;br /&gt;
The source does not establish whether these are valid nighttime observations or mistyped times. Changing the production observation window requires a separate decision. No current GROOM_SCANS row occurs exactly at 04:00 or 20:00.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
   gs.gs_fol_b_focal_animid,&lt;br /&gt;
   gs.gs_time,&lt;br /&gt;
   gs.gs_b_chimp1_animid,&lt;br /&gt;
   gs.gs_b_chimp2_animid,&lt;br /&gt;
   gs.gs_direction,&lt;br /&gt;
   gs.gs_extracted_by,&lt;br /&gt;
   gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE gs.gs_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
   OR gs.gs_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
      gs.gs_fol_b_focal_animid,&lt;br /&gt;
      gs.gs_time,&lt;br /&gt;
      gs.gs_b_chimp1_animid,&lt;br /&gt;
      gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the 61 affected rows from the production B-record groom-scan loader. They do not overlap Problems #146, #147, or #148. Investigators must confirm the intended times or decide separately whether the production observation window should change. Keep the current production constraints intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#150) GROOM_SCANS unknown direction cannot use the production UNKPair role ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Thirty-one GROOM_SCANS rows have &amp;lt;code&amp;gt;GS_direction = &amp;#039;U&amp;#039;&amp;lt;/code&amp;gt;. The loader maps this source code to paired &amp;lt;code&amp;gt;UNKPair&amp;lt;/code&amp;gt; roles, which represent a dyadic interaction whose direction is unknown. The current production &amp;lt;code&amp;gt;roles_func&amp;lt;/code&amp;gt; and EVENTS documentation permit Actor, Actee, and Mutual roles for GSCAN events but reject &amp;lt;code&amp;gt;UNKPair&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The first runtime failure is the 1982-02-07 ST scan at 11:55 with participants FF and GB. Its source comment says &amp;quot;NO DIRECTION INDICATED ON TIKI OR ON B RECORDS&amp;quot;. None of the 31 rows overlaps Problems #146 through #149.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE gs.gs_direction = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the 31 source rows while retaining the current production role restriction. The investigator does not want structural schema changes in the B-record loader commit. Flag permitting paired &amp;lt;code&amp;gt;UNKPair&amp;lt;/code&amp;gt; roles for GSCAN events, with the corresponding EVENTS documentation change, as a structural problem for the next iteration.&lt;br /&gt;
&lt;br /&gt;
Do not convert &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; to Mutual because unknown direction does not establish symmetric grooming.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#151) GROOM_SCANS participants are absent from BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Two hundred twenty-five GROOM_SCANS rows have a grooming participant absent from BIOGRAPHY: 33 affected values are in &amp;lt;code&amp;gt;GS_B_chimp1_AnimId&amp;lt;/code&amp;gt; and 192 are in &amp;lt;code&amp;gt;GS_B_chimp2_AnimId&amp;lt;/code&amp;gt;. Production ROLES.Participant must reference BIOGRAPHY_DATA.AnimID. The first runtime failure is the 1982-03-09 FF scan at 15:45, whose second participant is TP.&lt;br /&gt;
&lt;br /&gt;
The source does not identify a production animal that can safely replace an absent participant.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       missing_participant.source_column,&lt;br /&gt;
       missing_participant.animid,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GS_B_chimp1_AnimId&amp;#039;, gs.gs_b_chimp1_animid),&lt;br /&gt;
                (&amp;#039;GS_B_chimp2_AnimId&amp;#039;, gs.gs_b_chimp2_animid)&lt;br /&gt;
       ) AS missing_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography&lt;br /&gt;
          WHERE biography.b_animid = missing_participant.animid)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid,&lt;br /&gt;
          missing_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the 225 affected rows from the production B-record groom-scan loader. They do not overlap Problems #146 through #150. Investigators must identify the intended animals and correct the Access data. Keep the production foreign key to BIOGRAPHY_DATA intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#152) GROOM_SCANS direction values cannot be mapped to production roles ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After the Problem #136 whitespace normalization, 48 GROOM_SCANS rows have a direction other than G, R, M, or U: 43 are NULL, one is F, one is H, and three are T. The first runtime failure that reached the direction mapping has H on the 1989-06-22 GB scan at 08:15.&lt;br /&gt;
&lt;br /&gt;
The source does not establish which participant groomed the other, whether grooming was mutual, or whether direction was unknown. The loader cannot choose production roles safely.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE COALESCE(gs.gs_direction, &amp;#039;&amp;#039;) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the 48 affected rows from the production B-record groom-scan loader. Three overlap Problems #146 through #151, so this exclusion adds 45 rows to the excluded union. Investigators must determine the intended direction and correct the Access data. Do not infer roles from participant order or map missing direction to Mutual or unknown without a documented decision.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#153) GROOM_SCANS records the same animal as both participants ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Twelve GROOM_SCANS rows identify the same animal in &amp;lt;code&amp;gt;GS_B_chimp1_AnimId&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;GS_B_chimp2_AnimId&amp;lt;/code&amp;gt;. Production represents groom scans as dyadic events and requires each ROLES participant to be unique within an event. The first runtime failure is the 2015-03-16 NUR/NUR scan at 13:00, whose source comment says &amp;quot;Same ID for chimp1 and chimp2&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE gs.gs_b_chimp1_animid = gs.gs_b_chimp2_animid&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the 12 affected rows from the production B-record groom-scan loader. They do not overlap Problems #146 through #152. Investigators must identify the intended first or second participant, or decide separately how self-grooming should be represented. Keep the production requirement that a participant occur at most once in an event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#154) GROOM_SCANS records have no observation time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Eight GROOM_SCANS rows have NULL &amp;lt;code&amp;gt;GS_time&amp;lt;/code&amp;gt; values. Production EVENTS requires a time, and all eight source comments say &amp;quot;no time&amp;quot;. None of the rows overlaps Problems #146 through #153.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE gs.gs_time IS NULL&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the eight affected rows from the production B-record groom-scan loader. Investigators must recover the observation times from another source or document a separate policy for missing event times. Do not invent a time from the scan date, record order, or neighboring observations.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#155) FERTILITY boundary-type codes need production support rows ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The 168 &amp;lt;code&amp;gt;clean.fertility&amp;lt;/code&amp;gt; rows use four &amp;lt;code&amp;gt;StartType&amp;lt;/code&amp;gt; values and four &amp;lt;code&amp;gt;StopType&amp;lt;/code&amp;gt;&lt;br /&gt;
values. Production FERTILITY requires these values to reference&lt;br /&gt;
FERTILITY_STARTS and FERTILITY_STOPS, but both support tables are empty.&lt;br /&gt;
&lt;br /&gt;
The fertility domains exactly match the BIOGRAPHY_DATA boundary-type domains:&lt;br /&gt;
&amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;C&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;I&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;O&amp;lt;/code&amp;gt; match ENTRYTYPES, while &amp;lt;code&amp;gt;D&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;E&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;O&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;P&amp;lt;/code&amp;gt; match&lt;br /&gt;
DEPARTTYPES.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;#039;StartType&amp;#039; AS source_column,&lt;br /&gt;
       fertility.starttype AS code,&lt;br /&gt;
       COUNT(*) AS rows&lt;br /&gt;
  FROM clean.fertility AS fertility&lt;br /&gt;
  GROUP BY fertility.starttype&lt;br /&gt;
UNION ALL&lt;br /&gt;
SELECT &amp;#039;StopType&amp;#039;,&lt;br /&gt;
       fertility.stoptype,&lt;br /&gt;
       COUNT(*)&lt;br /&gt;
  FROM clean.fertility AS fertility&lt;br /&gt;
  GROUP BY fertility.stoptype&lt;br /&gt;
ORDER BY source_column, code;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed 2026-09-13 source contains StartType &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt; (70 rows), &amp;lt;code&amp;gt;C&amp;lt;/code&amp;gt; (56),&lt;br /&gt;
&amp;lt;code&amp;gt;I&amp;lt;/code&amp;gt; (34), and &amp;lt;code&amp;gt;O&amp;lt;/code&amp;gt; (8), and StopType &amp;lt;code&amp;gt;D&amp;lt;/code&amp;gt; (74), &amp;lt;code&amp;gt;E&amp;lt;/code&amp;gt; (1), &amp;lt;code&amp;gt;O&amp;lt;/code&amp;gt; (75), and &amp;lt;code&amp;gt;P&amp;lt;/code&amp;gt;&lt;br /&gt;
(18).&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The project approved using the corresponding ENTRYTYPES and DEPARTTYPES&lt;br /&gt;
meanings for fertility boundary types. Populate FERTILITY_STARTS with &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;&lt;br /&gt;
(Birth), &amp;lt;code&amp;gt;C&amp;lt;/code&amp;gt; (Start of confirmed AnimID), &amp;lt;code&amp;gt;I&amp;lt;/code&amp;gt; (Immigration), and &amp;lt;code&amp;gt;O&amp;lt;/code&amp;gt;&lt;br /&gt;
(Initiation of close observation). Populate FERTILITY_STOPS with &amp;lt;code&amp;gt;D&amp;lt;/code&amp;gt; (Death),&lt;br /&gt;
&amp;lt;code&amp;gt;E&amp;lt;/code&amp;gt; (Emigration), &amp;lt;code&amp;gt;O&amp;lt;/code&amp;gt; (End of observation; Present in the most recent census),&lt;br /&gt;
and &amp;lt;code&amp;gt;P&amp;lt;/code&amp;gt; (Permanent disappearance).&lt;br /&gt;
&lt;br /&gt;
This is a support-table decision and excludes no source rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#156) FERTILITY source intervals require ordered dates ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production FERTILITY does not have a table constraint requiring &amp;lt;code&amp;gt;StartDate &amp;lt;=&lt;br /&gt;
StopDate&amp;lt;/code&amp;gt;. Loading a reversed source interval would therefore preserve an&lt;br /&gt;
invalid period without necessarily producing a database error.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fertility.studyid,&lt;br /&gt;
       fertility.animalid,&lt;br /&gt;
       fertility.startdate,&lt;br /&gt;
       fertility.starttype,&lt;br /&gt;
       fertility.stopdate,&lt;br /&gt;
       fertility.stoptype&lt;br /&gt;
  FROM clean.fertility AS fertility&lt;br /&gt;
  WHERE fertility.startdate &amp;gt; fertility.stopdate&lt;br /&gt;
  ORDER BY fertility.studyid,&lt;br /&gt;
           fertility.animalid,&lt;br /&gt;
           fertility.startdate,&lt;br /&gt;
           fertility.stopdate;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed 2026-09-13 source contains zero reversed intervals.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Enforce &amp;lt;code&amp;gt;StartDate &amp;lt;= StopDate&amp;lt;/code&amp;gt; in the read-only fertility conversion sanity&lt;br /&gt;
gate. Do not add a production table constraint as part of this conversion.&lt;br /&gt;
Abort before writing any production rows if a future source snapshot contains&lt;br /&gt;
a reversed interval.&lt;br /&gt;
&lt;br /&gt;
This decision excludes no source rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#157) Selected FERTILITY periods cross BIOGRAPHY_DATA boundaries ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Nine &amp;lt;code&amp;gt;clean.fertility&amp;lt;/code&amp;gt; rows extend beyond one or more BIOGRAPHY_DATA date&lt;br /&gt;
boundaries. Production does not require fertility periods to remain within&lt;br /&gt;
BirthDate, EntryDate, and DepartDate, so the database does not establish&lt;br /&gt;
whether these differences are valid history or source errors.&lt;br /&gt;
&lt;br /&gt;
Investigators reviewed the nine rows and selected six for temporary exclusion:&lt;br /&gt;
AR, BH, GLIB1, MAM, NV, and RAF. The remaining boundary differences for SAF,&lt;br /&gt;
SHO, and SN1 are approved for conversion and are not excluded.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader; all six source fields identify each approved exclusion&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fertility.studyid,&lt;br /&gt;
       fertility.animalid,&lt;br /&gt;
       fertility.startdate,&lt;br /&gt;
       fertility.starttype,&lt;br /&gt;
       fertility.stopdate,&lt;br /&gt;
       fertility.stoptype&lt;br /&gt;
  FROM clean.fertility AS fertility&lt;br /&gt;
  WHERE (fertility.studyid,&lt;br /&gt;
         fertility.animalid,&lt;br /&gt;
         fertility.startdate,&lt;br /&gt;
         fertility.starttype,&lt;br /&gt;
         fertility.stopdate,&lt;br /&gt;
         fertility.stoptype)&lt;br /&gt;
        IN ((4, &amp;#039;AR&amp;#039;, DATE &amp;#039;1985-07-11&amp;#039;, &amp;#039;B&amp;#039;, DATE &amp;#039;1987-07-19&amp;#039;, &amp;#039;D&amp;#039;),&lt;br /&gt;
            (4, &amp;#039;BH&amp;#039;, DATE &amp;#039;1971-07-05&amp;#039;, &amp;#039;B&amp;#039;, DATE &amp;#039;1971-09-20&amp;#039;, &amp;#039;D&amp;#039;),&lt;br /&gt;
            (4, &amp;#039;GLIB1&amp;#039;, DATE &amp;#039;2011-07-11&amp;#039;, &amp;#039;B&amp;#039;, DATE &amp;#039;2012-01-21&amp;#039;, &amp;#039;D&amp;#039;),&lt;br /&gt;
            (4, &amp;#039;MAM&amp;#039;, DATE &amp;#039;2004-02-12&amp;#039;, &amp;#039;B&amp;#039;, DATE &amp;#039;2011-12-07&amp;#039;, &amp;#039;D&amp;#039;),&lt;br /&gt;
            (4, &amp;#039;NV&amp;#039;, DATE &amp;#039;1965-09-15&amp;#039;, &amp;#039;I&amp;#039;, DATE &amp;#039;1975-03-12&amp;#039;, &amp;#039;D&amp;#039;),&lt;br /&gt;
            (4, &amp;#039;RAF&amp;#039;, DATE &amp;#039;1986-10-21&amp;#039;, &amp;#039;C&amp;#039;, DATE &amp;#039;1996-03-27&amp;#039;, &amp;#039;D&amp;#039;))&lt;br /&gt;
  ORDER BY fertility.animalid,&lt;br /&gt;
           fertility.startdate;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The query returns exactly six rows. Because this is a new FERTILITY loader,&lt;br /&gt;
none overlaps an earlier source-row exclusion. The distinct exclusion union is&lt;br /&gt;
six rows, leaving 162 eligible rows from the refreshed 168-row source.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the six identified rows from the production FERTILITY&lt;br /&gt;
loader. Keep the full six-column predicate adjacent to its diagnostic query so&lt;br /&gt;
the exclusion cannot broaden silently. Investigators must correct the Access&lt;br /&gt;
fertility or biography dates, or approve a later interpretation that permits&lt;br /&gt;
these periods, before the rows are restored.&lt;br /&gt;
&lt;br /&gt;
Do not exclude SAF, SHO, or SN1. Their boundary differences were reviewed and&lt;br /&gt;
approved for conversion.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#158) FERTILITY StudyId 4 has no production STUDIES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All 168 &amp;lt;code&amp;gt;clean.fertility&amp;lt;/code&amp;gt; rows have integer &amp;lt;code&amp;gt;StudyId&amp;lt;/code&amp;gt; 4. Production&lt;br /&gt;
FERTILITY.Study is text and must reference STUDIES, but the follow-derived&lt;br /&gt;
support data contains no Study code &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt;. The Access snapshot contains no study&lt;br /&gt;
lookup table, description, foreign key, or imported metadata that gives StudyId&lt;br /&gt;
4 a textual name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fertility.studyid,&lt;br /&gt;
       COUNT(*) AS rows&lt;br /&gt;
  FROM clean.fertility AS fertility&lt;br /&gt;
  GROUP BY fertility.studyid&lt;br /&gt;
  ORDER BY fertility.studyid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed 2026-09-13 source returns only StudyId 4, with 168 rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Preserve the source value without inventing unsupported semantics. Add Study&lt;br /&gt;
code &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt; to STUDIES with description &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt;, and map &amp;lt;code&amp;gt;clean.fertility.studyid&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;FERTILITY.Study&amp;lt;/code&amp;gt; using &amp;lt;code&amp;gt;studyid::text&amp;lt;/code&amp;gt; at the load boundary.&lt;br /&gt;
&lt;br /&gt;
This support and representation decision excludes no source rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#159) FOOD_LOOKUP status flags need explicit conversion semantics ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;clean.food_lookup&amp;lt;/code&amp;gt; contains &amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt; and&lt;br /&gt;
&amp;lt;code&amp;gt;FL_unverified&amp;lt;/code&amp;gt; flags, but the mung support and event loaders ignored&lt;br /&gt;
both.  Loading all lookup rows would admit names explicitly marked for&lt;br /&gt;
exclusion; treating unverified names the same way would discard data without&lt;br /&gt;
investigator approval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;The following is a dated diagnostic. Rerun it against the refreshed clean&lt;br /&gt;
schema and record exact referring rows before implementing exclusions.&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH normalized_lookup AS (&lt;br /&gt;
  SELECT UPPER(BTRIM(fl_local_food_name)) AS foodname,&lt;br /&gt;
         fl_exclude,&lt;br /&gt;
         &amp;quot;FL_unverified&amp;quot; AS fl_unverified&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
),&lt;br /&gt;
used AS (&lt;br /&gt;
  SELECT &amp;#039;seq1&amp;#039; AS slot,&lt;br /&gt;
         fb_fol_date,&lt;br /&gt;
         fb_fol_b_focal_animid,&lt;br /&gt;
         fb_begin_feed_time,&lt;br /&gt;
         fb_end_feed_time,&lt;br /&gt;
         UPPER(BTRIM(fb_fl_local_food_name)) AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;seq2&amp;#039;,&lt;br /&gt;
         fb_fol_date,&lt;br /&gt;
         fb_fol_b_focal_animid,&lt;br /&gt;
         fb_begin_feed_time,&lt;br /&gt;
         fb_end_feed_time,&lt;br /&gt;
         UPPER(BTRIM(fb_local_food_name2))&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    WHERE NULLIF(BTRIM(fb_local_food_name2), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT used.*,&lt;br /&gt;
       normalized_lookup.fl_exclude,&lt;br /&gt;
       normalized_lookup.fl_unverified&lt;br /&gt;
  FROM used&lt;br /&gt;
  JOIN normalized_lookup USING (foodname)&lt;br /&gt;
  WHERE normalized_lookup.fl_exclude IS TRUE&lt;br /&gt;
     OR normalized_lookup.fl_unverified IS TRUE&lt;br /&gt;
  ORDER BY used.fb_fol_date,&lt;br /&gt;
           used.fb_fol_b_focal_animid,&lt;br /&gt;
           used.fb_begin_feed_time,&lt;br /&gt;
           used.slot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-13 snapshot had 18 excluded lookup rows and 359 unverified lookup&lt;br /&gt;
rows.  Excluded names were referenced by 538 primary and 24 secondary bout&lt;br /&gt;
items, 562 occurrences total.  Unverified names were referenced by 728 bout&lt;br /&gt;
items.  These sets may overlap and must be profiled separately.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;FL_exclude = true&amp;lt;/code&amp;gt; disqualifies the lookup entry.  Do not load that&lt;br /&gt;
name into &amp;lt;code&amp;gt;FOOD_NAMES&amp;lt;/code&amp;gt;, and exactly exclude every referring food-bout&lt;br /&gt;
item after approved case and edge-whitespace normalization.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;FL_unverified = true&amp;lt;/code&amp;gt; does not disqualify a lookup entry or&lt;br /&gt;
referring bout.  Preserve the flag in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt;, document that it has no&lt;br /&gt;
production destination, and load the otherwise eligible name and bout.&lt;br /&gt;
&lt;br /&gt;
The sanity gate must prove that no excluded name enters production and that no&lt;br /&gt;
unverified name is omitted merely because it is unverified.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#160) Converted food events require a consumer ROLES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The mung food loader inserted &amp;lt;code&amp;gt;EVENTS&amp;lt;/code&amp;gt; and&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_EVENTS&amp;lt;/code&amp;gt; rows but inserted no &amp;lt;code&amp;gt;ROLES&amp;lt;/code&amp;gt; row.  Production&lt;br /&gt;
documentation says the food event&amp;#039;s role identifies the individual consuming&lt;br /&gt;
the food.  The roles trigger permits at most one role and requires that&lt;br /&gt;
participant to equal the related watch focal.  Although absence currently&lt;br /&gt;
produces a warning rather than a hard constraint failure, omitting the consumer&lt;br /&gt;
would make the conversion incomplete.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Run after a rollback-only or disposable-database load.&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT events.eid,&lt;br /&gt;
       watches.wid,&lt;br /&gt;
       watches.animid AS focal,&lt;br /&gt;
       COUNT(roles.pid) AS role_rows,&lt;br /&gt;
       MIN(roles.role) AS role,&lt;br /&gt;
       MIN(roles.participant) AS participant&lt;br /&gt;
  FROM sokwedb.events&lt;br /&gt;
  JOIN sokwedb.watches&lt;br /&gt;
    ON watches.wid = events.wid&lt;br /&gt;
  LEFT JOIN sokwedb.roles&lt;br /&gt;
    ON roles.eid = events.eid&lt;br /&gt;
  WHERE events.behavior = &amp;#039;FOOD&amp;#039;&lt;br /&gt;
  GROUP BY events.eid,&lt;br /&gt;
           watches.wid,&lt;br /&gt;
           watches.animid&lt;br /&gt;
  HAVING COUNT(roles.pid) &amp;lt;&amp;gt; 1&lt;br /&gt;
      OR MIN(roles.role) &amp;lt;&amp;gt; &amp;#039;Solo&amp;#039;&lt;br /&gt;
      OR MIN(roles.participant) &amp;lt;&amp;gt; watches.animid&lt;br /&gt;
  ORDER BY events.eid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For every eligible food bout, insert exactly one &amp;lt;code&amp;gt;ROLES&amp;lt;/code&amp;gt; row related&lt;br /&gt;
to the newly created event.  Use role code &amp;lt;code&amp;gt;Solo&amp;lt;/code&amp;gt; and the selected B&lt;br /&gt;
watch&amp;#039;s exact &amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt; as Participant.  Capture the generated EID;&lt;br /&gt;
do not infer event identity from sequence state or source ordering.&lt;br /&gt;
&lt;br /&gt;
This decision excludes no source rows.  The loader and parity checks must prove&lt;br /&gt;
one food event, one Solo consumer role, and one or two contiguous food-detail&lt;br /&gt;
rows per eligible source bout.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#161) FOOD_VARIATIONS conversion is deferred ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;clean.food_variations_lookup&amp;lt;/code&amp;gt; has a corresponding production&lt;br /&gt;
&amp;lt;code&amp;gt;housekeeping.FOOD_VARIATIONS&amp;lt;/code&amp;gt; table, but the mung branch never&lt;br /&gt;
loaded it.  The current food-event conversion requires decisions for missing&lt;br /&gt;
and unresolved LocalName values that are independent of loading food bouts and&lt;br /&gt;
their two support tables.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;This is a dated scope diagnostic, not an exclusion query.&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) AS source_rows,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE NULLIF(BTRIM(fvl_food_spelling_variant), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
       ) AS missing_variants,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE NULLIF(BTRIM(fvl_fl_local_food_name), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
       ) AS missing_local_names&lt;br /&gt;
  FROM clean.food_variations_lookup;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-13 snapshot contained 1,575 rows, no blank variants, 224 blank or&lt;br /&gt;
NULL local names, and 38 additional nonblank local names without a normalized&lt;br /&gt;
match in &amp;lt;code&amp;gt;clean.food_lookup&amp;lt;/code&amp;gt;.  Rerun and extend the query in the&lt;br /&gt;
future variations conversion; do not reuse these counts as exclusions.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Explicitly defer &amp;lt;code&amp;gt;clean.food_variations_lookup&amp;lt;/code&amp;gt; to a separate future&lt;br /&gt;
conversion.  The present food conversion must not insert&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_VARIATIONS&amp;lt;/code&amp;gt; rows, must not depend on that table to normalize&lt;br /&gt;
bout names, and must not claim variation parity.&lt;br /&gt;
&lt;br /&gt;
Preserve this issue and the source table for the future conversion.  The&lt;br /&gt;
deferral excludes no food-bout rows and does not authorize dropping variation&lt;br /&gt;
data.&lt;br /&gt;
&lt;br /&gt;
== (#162) Eligible FOOD_LOOKUP rows lack generalized scientific names ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The approved &amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt; source is&lt;br /&gt;
&amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt;, but 360 eligible lookup rows have a blank or&lt;br /&gt;
NULL value.  All 360 names are used by food bouts, in 812 primary and 111&lt;br /&gt;
secondary items.  Of these names, 359 are marked &amp;lt;code&amp;gt;FL_unverified&amp;lt;/code&amp;gt; and&lt;br /&gt;
must not be disqualified under Problem #159.  The target description is NOT&lt;br /&gt;
NULL, nonempty, and case-insensitively unique.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-13: use&lt;br /&gt;
&amp;lt;code&amp;gt;Unknown -- &amp;amp;lt;local name&amp;amp;gt;&amp;lt;/code&amp;gt; for an eligible blank generalized name.&lt;br /&gt;
This stable value identifies the exact source name and does not consult&lt;br /&gt;
&amp;lt;code&amp;gt;fl_sci_food_name&amp;lt;/code&amp;gt;.  Nonblank descriptions continue to follow&lt;br /&gt;
Problem #87.  This mapping excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#163) FOOD_BOUT duration disagrees with elapsed event time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The refreshed source has 1,136 rows where &amp;lt;code&amp;gt;fb_duration&amp;lt;/code&amp;gt; differs from&lt;br /&gt;
the exact minutes between begin and end.  Production stores the original begin&lt;br /&gt;
and end times but has no independent duration column.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not alter event times to reproduce &amp;lt;code&amp;gt;fb_duration&amp;lt;/code&amp;gt;.  Retain the&lt;br /&gt;
field in &amp;lt;code&amp;gt;clean.food_bout&amp;lt;/code&amp;gt;, emit a sanity warning with the refreshed&lt;br /&gt;
count, and document that it has no production destination.  This issue causes&lt;br /&gt;
no exclusions.&lt;br /&gt;
&lt;br /&gt;
== (#164) FOOD_BOUT update metadata has no production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The refreshed source has 17,702 rows with non-NULL &amp;lt;code&amp;gt;fb_update&amp;lt;/code&amp;gt;.&lt;br /&gt;
None of the food target tables has a corresponding update-metadata column.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Retain &amp;lt;code&amp;gt;fb_update&amp;lt;/code&amp;gt; in &amp;lt;code&amp;gt;clean.food_bout&amp;lt;/code&amp;gt;, emit a sanity&lt;br /&gt;
warning with the refreshed count, and do not overload event notes or another&lt;br /&gt;
production field.  This issue causes no exclusions.&lt;br /&gt;
&lt;br /&gt;
== * (#165) ATTENDANCE contains exact duplicate rows ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The source table has no primary key and contains values that are identical in&lt;br /&gt;
all 17 columns.  Loading every copy would create indistinguishable production&lt;br /&gt;
events and repeated sequence values.  Retaining an arbitrary copy would not be&lt;br /&gt;
an exact source-row policy because no source value distinguishes the copies.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Run against the refreshed &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  The grouping explicitly&lt;br /&gt;
compares all 17 typed source columns.&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*,&lt;br /&gt;
       COUNT(*) AS copies&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  GROUP BY attendance.a_date,&lt;br /&gt;
           attendance.a_cl_community_id,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num,&lt;br /&gt;
           attendance.a_type_of_cycle,&lt;br /&gt;
           attendance.a_observer1_id,&lt;br /&gt;
           attendance.a_observer2_id,&lt;br /&gt;
           attendance.a_bananas_given,&lt;br /&gt;
           attendance.a_degree_of_arrival,&lt;br /&gt;
           attendance.a_degree_of_departure,&lt;br /&gt;
           attendance.a_time_start,&lt;br /&gt;
           attendance.a_time_end,&lt;br /&gt;
           attendance.a_duration_of_obs,&lt;br /&gt;
           attendance.day,&lt;br /&gt;
           attendance.mo,&lt;br /&gt;
           attendance.yr,&lt;br /&gt;
           attendance.cycle_old&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num,&lt;br /&gt;
           attendance.a_time_start;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot contained 207 duplicate values, each occurring twice:&lt;br /&gt;
414 source rows total.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exclude every copy of each exact&lt;br /&gt;
duplicate value.  Do not use &amp;lt;code&amp;gt;DISTINCT&amp;lt;/code&amp;gt; to retain one arbitrary copy.&lt;br /&gt;
The shared attendance projection and loader must include this query beside the&lt;br /&gt;
Problem #165 predicate and prove that every row whose complete typed value has&lt;br /&gt;
multiplicity greater than one is excluded.&lt;br /&gt;
&lt;br /&gt;
Implemented in &amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt; as&lt;br /&gt;
&amp;lt;code&amp;gt;duplicate_count &amp;amp;gt; 1&amp;lt;/code&amp;gt;, partitioned over all 17 columns.  On&lt;br /&gt;
2026-09-15 the investigator approved semantic validation: the dated 207 values&lt;br /&gt;
and 414 rows are audit evidence, not frozen acceptance criteria.  Sanity&lt;br /&gt;
asserts that every refreshed duplicate copy is classified and excluded.&lt;br /&gt;
Remove only when exact duplicates leave the source.&lt;br /&gt;
&lt;br /&gt;
== * (#166) ATTENDANCE sequence values are not contiguous per animal and date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production attendance sequence values are expected to be exactly&lt;br /&gt;
&amp;lt;code&amp;gt;1..N&amp;lt;/code&amp;gt; for each animal/date.  Exact duplicates and other targeted&lt;br /&gt;
row exclusions can themselves create duplicate values or gaps, so sequence&lt;br /&gt;
validity must be checked after all other attendance exclusions.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;This query expresses the approved staging against the 2026-09-14 schema.&lt;br /&gt;
Keep it synchronized with Problems #165 and #168 through #178 when predicates&lt;br /&gt;
change.&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source AS (&lt;br /&gt;
  SELECT row_number() OVER () AS source_id,&lt;br /&gt;
         attendance.*,&lt;br /&gt;
         COUNT(*) OVER (&lt;br /&gt;
           PARTITION BY attendance.a_date,&lt;br /&gt;
                        attendance.a_cl_community_id,&lt;br /&gt;
                        attendance.a_b_animid,&lt;br /&gt;
                        attendance.a_seq_num,&lt;br /&gt;
                        attendance.a_type_of_cycle,&lt;br /&gt;
                        attendance.a_observer1_id,&lt;br /&gt;
                        attendance.a_observer2_id,&lt;br /&gt;
                        attendance.a_bananas_given,&lt;br /&gt;
                        attendance.a_degree_of_arrival,&lt;br /&gt;
                        attendance.a_degree_of_departure,&lt;br /&gt;
                        attendance.a_time_start,&lt;br /&gt;
                        attendance.a_time_end,&lt;br /&gt;
                        attendance.a_duration_of_obs,&lt;br /&gt;
                        attendance.day,&lt;br /&gt;
                        attendance.mo,&lt;br /&gt;
                        attendance.yr,&lt;br /&gt;
                        attendance.cycle_old&lt;br /&gt;
         ) AS duplicate_count&lt;br /&gt;
    FROM clean.attendance AS attendance&lt;br /&gt;
),&lt;br /&gt;
independent_eligible AS (&lt;br /&gt;
  SELECT source.*&lt;br /&gt;
    FROM source&lt;br /&gt;
      JOIN sokwedb.biography_data AS biography&lt;br /&gt;
        ON biography.animid = source.a_b_animid::text&lt;br /&gt;
    WHERE source.duplicate_count = 1&lt;br /&gt;
      AND source.a_date BETWEEN biography.entrydate AND biography.departdate&lt;br /&gt;
      AND COALESCE(BTRIM(source.a_observer1_id::text), &amp;#039;NONE&amp;#039;)&lt;br /&gt;
            &amp;amp;lt;&amp;amp;gt; COALESCE(BTRIM(source.a_observer2_id::text), &amp;#039;NONE&amp;#039;)&lt;br /&gt;
      AND BTRIM(source.a_type_of_cycle::text)&lt;br /&gt;
            NOT IN (&amp;#039;-&amp;#039;, &amp;#039;0.30&amp;#039;, &amp;#039;0.33&amp;#039;)&lt;br /&gt;
      AND source.cycle_old IS NOT NULL&lt;br /&gt;
      AND source.cycle_old::text = BTRIM(source.cycle_old::text)&lt;br /&gt;
      AND (source.a_bananas_given IS NULL&lt;br /&gt;
           OR source.a_bananas_given = TRUNC(source.a_bananas_given))&lt;br /&gt;
      AND source.a_time_start IS NOT NULL&lt;br /&gt;
      AND source.a_time_end IS NOT NULL&lt;br /&gt;
      AND source.a_time_start &amp;amp;lt;= source.a_time_end&lt;br /&gt;
      AND source.a_duration_of_obs IS NOT DISTINCT FROM&lt;br /&gt;
            EXTRACT(EPOCH FROM&lt;br /&gt;
                    (source.a_time_end - source.a_time_start)) / 60&lt;br /&gt;
      AND source.day IS NOT DISTINCT FROM&lt;br /&gt;
        EXTRACT(DAY FROM source.a_date)::integer&lt;br /&gt;
      AND source.mo IS NOT DISTINCT FROM&lt;br /&gt;
        EXTRACT(MONTH FROM source.a_date)::integer&lt;br /&gt;
      AND source.yr IS NOT DISTINCT FROM&lt;br /&gt;
        EXTRACT(YEAR FROM source.a_date)::integer&lt;br /&gt;
),&lt;br /&gt;
overlap_rows AS (&lt;br /&gt;
  SELECT DISTINCT first_row.source_id&lt;br /&gt;
    FROM independent_eligible AS first_row&lt;br /&gt;
      JOIN independent_eligible AS second_row&lt;br /&gt;
        ON second_row.source_id &amp;amp;lt;&amp;amp;gt; first_row.source_id&lt;br /&gt;
       AND second_row.a_date = first_row.a_date&lt;br /&gt;
       AND second_row.a_b_animid = first_row.a_b_animid&lt;br /&gt;
       AND second_row.a_time_start &amp;amp;lt;= first_row.a_time_end&lt;br /&gt;
       AND second_row.a_time_end &amp;amp;gt;= first_row.a_time_start&lt;br /&gt;
),&lt;br /&gt;
eligible_before_seq AS (&lt;br /&gt;
  SELECT independent_eligible.*&lt;br /&gt;
    FROM independent_eligible&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
      SELECT 1&lt;br /&gt;
        FROM overlap_rows&lt;br /&gt;
          WHERE overlap_rows.source_id = independent_eligible.source_id)&lt;br /&gt;
),&lt;br /&gt;
bad_groups AS (&lt;br /&gt;
  SELECT a_date,&lt;br /&gt;
         a_b_animid&lt;br /&gt;
    FROM eligible_before_seq&lt;br /&gt;
    GROUP BY a_date,&lt;br /&gt;
             a_b_animid&lt;br /&gt;
    HAVING MIN(a_seq_num) IS NULL&lt;br /&gt;
       OR MIN(a_seq_num) &amp;amp;lt;&amp;amp;gt; 1&lt;br /&gt;
       OR MAX(a_seq_num) IS DISTINCT FROM COUNT(*)&lt;br /&gt;
       OR COUNT(DISTINCT a_seq_num) &amp;amp;lt;&amp;amp;gt; COUNT(*)&lt;br /&gt;
)&lt;br /&gt;
SELECT eligible_before_seq.*&lt;br /&gt;
  FROM eligible_before_seq&lt;br /&gt;
    JOIN bad_groups&lt;br /&gt;
      USING (a_date, a_b_animid)&lt;br /&gt;
  ORDER BY eligible_before_seq.a_date,&lt;br /&gt;
           eligible_before_seq.a_b_animid,&lt;br /&gt;
           eligible_before_seq.a_seq_num,&lt;br /&gt;
           eligible_before_seq.a_time_start;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed projection contained 453 invalid animal/date groups and 910 rows&lt;br /&gt;
after the other approved exclusions and the Problem #171 mapping.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exclude every row in each invalid&lt;br /&gt;
animal/date group.  Do not renumber, retain a partial group, or select rows by&lt;br /&gt;
order.  Compute Problem #166 last, after Problems #165, #168 through #170, and&lt;br /&gt;
#172 through #178,&lt;br /&gt;
and include the complete identification query beside the loader predicate.&lt;br /&gt;
Sanity must prove every remaining group has non-NULL sequence values exactly&lt;br /&gt;
equal to &amp;lt;code&amp;gt;1..N&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Implemented last in &amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt; by&lt;br /&gt;
&amp;lt;code&amp;gt;attendance_bad_seq_groups&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;problem_166&amp;lt;/code&amp;gt;.  The&lt;br /&gt;
dated 453 groups and 910 rows are audit evidence.  Sanity proves every&lt;br /&gt;
refreshed invalid group is completely excluded and every eligible group is&lt;br /&gt;
exactly &amp;lt;code&amp;gt;1..N&amp;lt;/code&amp;gt;.  Remove only when all post-exclusion groups are&lt;br /&gt;
contiguous.&lt;br /&gt;
&lt;br /&gt;
== (#167) ATTENDANCE observer values require PEOPLE support rows ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The observer columns contain hundreds of case and edge-whitespace variants.&lt;br /&gt;
Observer codes must be matched by their standardized value so variants such as&lt;br /&gt;
&amp;lt;code&amp;gt;Mike&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;mike&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;Mike &amp;lt;/code&amp;gt; resolve to one&lt;br /&gt;
&amp;lt;code&amp;gt;PEOPLE.Person&amp;lt;/code&amp;gt;.  Standardized values absent from&lt;br /&gt;
&amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; require support rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH raw_observers AS (&lt;br /&gt;
  SELECT a_observer1_id::text AS raw_person&lt;br /&gt;
    FROM clean.attendance&lt;br /&gt;
    WHERE a_observer1_id IS NOT NULL&lt;br /&gt;
  UNION&lt;br /&gt;
  SELECT a_observer2_id::text&lt;br /&gt;
    FROM clean.attendance&lt;br /&gt;
    WHERE a_observer2_id IS NOT NULL&lt;br /&gt;
), standardized AS (&lt;br /&gt;
  SELECT raw_person,&lt;br /&gt;
         LOWER(NORMALIZE(BTRIM(raw_person))) AS standard_person&lt;br /&gt;
    FROM raw_observers&lt;br /&gt;
)&lt;br /&gt;
SELECT standardized.standard_person,&lt;br /&gt;
       ARRAY_AGG(standardized.raw_person&lt;br /&gt;
                 ORDER BY standardized.raw_person) AS raw_spellings,&lt;br /&gt;
       ARRAY_AGG(people.person ORDER BY people.person)&lt;br /&gt;
         FILTER (WHERE people.person IS NOT NULL) AS existing_codes&lt;br /&gt;
  FROM standardized&lt;br /&gt;
  LEFT JOIN codes.people&lt;br /&gt;
    ON LOWER(NORMALIZE(BTRIM(people.person))) =&lt;br /&gt;
         standardized.standard_person&lt;br /&gt;
  GROUP BY standardized.standard_person&lt;br /&gt;
  ORDER BY standardized.standard_person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed 2026-09-14 snapshot had 932 distinct raw observer spellings in&lt;br /&gt;
718 standardized classes.  Of those classes, 245 matched an existing&lt;br /&gt;
&amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; row and 473 required a new support row.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: standardize observer identifiers&lt;br /&gt;
with &amp;lt;code&amp;gt;LOWER(NORMALIZE(BTRIM(value)))&amp;lt;/code&amp;gt;.  Resolve a class to its existing&lt;br /&gt;
&amp;lt;code&amp;gt;PEOPLE.Person&amp;lt;/code&amp;gt; spelling when present.  For an unmatched class, choose&lt;br /&gt;
one deterministic trimmed source spelling, preferring mixed case and then&lt;br /&gt;
bytewise order, and add it as an active support row.  Use that spelling for&lt;br /&gt;
Person and Name and &amp;lt;code&amp;gt;Attendance observer code: &amp;amp;lt;code&amp;amp;gt;&amp;lt;/code&amp;gt; for&lt;br /&gt;
Description.  Map NULL to the existing &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; code.  This support&lt;br /&gt;
decision excludes no rows except where Problem #168 also applies.&lt;br /&gt;
&lt;br /&gt;
Implemented by &amp;lt;code&amp;gt;attendance_observer_codes&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt; and the ordered&lt;br /&gt;
&amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; insert in &amp;lt;code&amp;gt;conversion/load_attendance.sql&amp;lt;/code&amp;gt;.  Sanity&lt;br /&gt;
pins all 718 canonical codes using an explicitly byte-ordered, length-prefixed&lt;br /&gt;
serialization, all 932 raw-to-canonical mappings, and 473 missing support rows.&lt;br /&gt;
Remove only when every standardized observer value already resolves to one&lt;br /&gt;
active row.&lt;br /&gt;
&lt;br /&gt;
== * (#168) ATTENDANCE Recorder and Observer2 resolve to the same person code ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;ARRIVALS_A&amp;lt;/code&amp;gt; requires Recorder and Observer2 to differ.  Some source&lt;br /&gt;
rows contain observer spellings that become equal after standardizing case and&lt;br /&gt;
edge whitespace; rows with both observers NULL also become equal after the&lt;br /&gt;
approved NULL-to-&amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; mapping.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE LOWER(NORMALIZE(COALESCE(&lt;br /&gt;
          BTRIM(attendance.a_observer1_id::text), &amp;#039;NONE&amp;#039;))) =&lt;br /&gt;
        LOWER(NORMALIZE(COALESCE(&lt;br /&gt;
          BTRIM(attendance.a_observer2_id::text), &amp;#039;NONE&amp;#039;)))&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num,&lt;br /&gt;
           attendance.a_time_start;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned 1,544 rows: 1,484 with equal trimmed&lt;br /&gt;
non-NULL spellings, 39 additional case-equivalent pairs, and 21 with both&lt;br /&gt;
observers NULL.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Do not fabricate or discard an observer.  Include the query beside the&lt;br /&gt;
Problem #168 predicate and prove no eligible detail has equal observer codes.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_168&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 1,544 rows are&lt;br /&gt;
audit evidence; sanity proves that no refreshed eligible row has equal derived&lt;br /&gt;
observer codes.  Remove only when no derived Recorder equals Observer2.&lt;br /&gt;
&lt;br /&gt;
== * (#169) ATTENDANCE focal does not resolve to BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An attendance watch and its Solo role require an exact production AnimID.&lt;br /&gt;
Thousands of source rows use values absent from &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt;,&lt;br /&gt;
mostly &amp;lt;code&amp;gt;NYANI&amp;lt;/code&amp;gt;.  Case or whitespace repair is not approved.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
      FROM sokwedb.biography_data AS biography&lt;br /&gt;
      WHERE biography.animid = attendance.a_b_animid::text)&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num,&lt;br /&gt;
           attendance.a_time_start;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The exact-match query returned 6,877 rows on 2026-09-14.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Do not trim, case-fold, or map an unmatched focal to a sentinel.  Include&lt;br /&gt;
the query beside the Problem #169 predicate and prove every eligible focal&lt;br /&gt;
resolves exactly.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_169&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 6,877 rows are&lt;br /&gt;
audit evidence; sanity proves that every refreshed eligible focal resolves.&lt;br /&gt;
Remove only when every exact source focal resolves.&lt;br /&gt;
&lt;br /&gt;
== * (#170) ATTENDANCE focal is outside its study dates ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A role participant must be under study on the event date.  Some exact-matching&lt;br /&gt;
focals have attendance dates before EntryDate or after DepartDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*,&lt;br /&gt;
       biography.entrydate,&lt;br /&gt;
       biography.departdate&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
    JOIN sokwedb.biography_data AS biography&lt;br /&gt;
      ON biography.animid = attendance.a_b_animid::text&lt;br /&gt;
  WHERE attendance.a_date NOT BETWEEN&lt;br /&gt;
        biography.entrydate AND biography.departdate&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned six rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Include the query beside the Problem #170 predicate and prove every&lt;br /&gt;
eligible participant is under study on the event date.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_170&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated six rows are audit&lt;br /&gt;
evidence; sanity proves that every refreshed eligible focal is under study on&lt;br /&gt;
its attendance date.  Remove only when every exact focal is under study.&lt;br /&gt;
&lt;br /&gt;
== (#171) ATTENDANCE direction fields contain the -9 sentinel ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production degrees may be NULL or integers from 0 through 359.  The source uses&lt;br /&gt;
&amp;lt;code&amp;gt;-9&amp;lt;/code&amp;gt; in one or both direction fields to mean that the direction was&lt;br /&gt;
not seen.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE COALESCE(attendance.a_degree_of_arrival = -9, FALSE)&lt;br /&gt;
     OR COALESCE(attendance.a_degree_of_departure = -9, FALSE)&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned 7,228 rows.  Individual arrival and departure&lt;br /&gt;
counts overlap and must not be summed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Confirmed by the project investigators on 2026-09-15: translate each&lt;br /&gt;
&amp;lt;code&amp;gt;-9&amp;lt;/code&amp;gt; independently to NULL because it means the direction was not&lt;br /&gt;
seen.  Preserve existing NULL and values from 0 through 359 exactly.  Keep this&lt;br /&gt;
query beside the Problem #171 predicate to profile the mapped source condition,&lt;br /&gt;
but do not add Problem #171 to the exclusion array.  Prove every eligible&lt;br /&gt;
target degree equals &amp;lt;code&amp;gt;NULLIF(source_degree, -9)&amp;lt;/code&amp;gt; and satisfies the&lt;br /&gt;
target domain.&lt;br /&gt;
&lt;br /&gt;
Implemented as the profiled &amp;lt;code&amp;gt;problem_171&amp;lt;/code&amp;gt; condition and the&lt;br /&gt;
&amp;lt;code&amp;gt;NULLIF(..., -9)&amp;lt;/code&amp;gt; arrival/departure projection in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 7,228 rows are&lt;br /&gt;
audit evidence; no rows are excluded solely by this issue.  Sanity proves each&lt;br /&gt;
refreshed eligible target equals &amp;lt;code&amp;gt;NULLIF(source_degree, -9)&amp;lt;/code&amp;gt;.  Remove&lt;br /&gt;
the mapping only when neither source degree contains &amp;lt;code&amp;gt;-9&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#172) ATTENDANCE swelling values have no approved CYCLE_STATES mapping ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Most source swelling spellings have an exact lossless representation in&lt;br /&gt;
&amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt;.  The values &amp;lt;code&amp;gt;-&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.30&amp;lt;/code&amp;gt;, and&lt;br /&gt;
&amp;lt;code&amp;gt;0.33&amp;lt;/code&amp;gt; do not.  Mapping &amp;lt;code&amp;gt;-&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; or rounding&lt;br /&gt;
the numeric values would change source meaning.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE BTRIM(attendance.a_type_of_cycle::text)&lt;br /&gt;
        IN (&amp;#039;-&amp;#039;, &amp;#039;0.30&amp;#039;, &amp;#039;0.33&amp;#039;)&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned 3,238 rows: 3,221 with &amp;lt;code&amp;gt;-&amp;lt;/code&amp;gt;, two with&lt;br /&gt;
&amp;lt;code&amp;gt;0.30&amp;lt;/code&amp;gt;, and 15 with &amp;lt;code&amp;gt;0.33&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Do not map or round these values.  The explicit supported projection is:&lt;br /&gt;
&amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;0.00&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;;&lt;br /&gt;
&amp;lt;code&amp;gt;.25&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;;&lt;br /&gt;
&amp;lt;code&amp;gt;.5&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;0.50&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;;&lt;br /&gt;
&amp;lt;code&amp;gt;.75&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;;&lt;br /&gt;
&amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;1.00&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt;;&lt;br /&gt;
&amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;N/A&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;; and&lt;br /&gt;
&amp;lt;code&amp;gt;u&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;.  Include the query beside the&lt;br /&gt;
Problem #172 predicate and reject any refreshed spelling outside this closed&lt;br /&gt;
mapping.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_172&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 3,238 rows are&lt;br /&gt;
audit evidence; sanity rejects any refreshed eligible spelling outside the&lt;br /&gt;
closed exact mapping.  Remove only when every source spelling belongs to it.&lt;br /&gt;
&lt;br /&gt;
== * (#173) ATTENDANCE CycleOld is missing or has edge spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;ARRIVALS_A.CycleOld&amp;lt;/code&amp;gt; is required and may not contain only spaces.&lt;br /&gt;
The source has NULL values and edge-spaced values.  Trimming would alter the&lt;br /&gt;
initial digitization recorded by this field.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE attendance.cycle_old IS NULL&lt;br /&gt;
     OR attendance.cycle_old::text &amp;amp;lt;&amp;amp;gt;&lt;br /&gt;
        BTRIM(attendance.cycle_old::text)&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned 417 rows: 229 NULL and 188 edge-spaced values.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Preserve all eligible CycleOld text exactly; do not trim or invent a&lt;br /&gt;
missing value.  Include the query beside the Problem #173 predicate.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_173&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 417 rows are audit&lt;br /&gt;
evidence; sanity proves every refreshed eligible CycleOld is non-NULL and&lt;br /&gt;
edge-space free.  Remove only when every source value satisfies that rule.&lt;br /&gt;
&lt;br /&gt;
== * (#174) ATTENDANCE bananas contains a fractional value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;ARRIVALS_A.Bananas&amp;lt;/code&amp;gt; is an integer.  One source value is fractional,&lt;br /&gt;
and rounding or truncating it would alter the recorded value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE attendance.a_bananas_given IS NOT NULL&lt;br /&gt;
    AND attendance.a_bananas_given &amp;amp;lt;&amp;amp;gt;&lt;br /&gt;
        TRUNC(attendance.a_bananas_given)&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned one row with value &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude the returned row.&lt;br /&gt;
Do not round or truncate it.  Include the query beside the Problem #174&lt;br /&gt;
predicate and prove all eligible banana values are integral and in range.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_174&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated one-row result is&lt;br /&gt;
audit evidence; sanity proves all refreshed eligible banana values are&lt;br /&gt;
integral and in range.  Remove only when all source values are integral.&lt;br /&gt;
&lt;br /&gt;
== * (#175) ATTENDANCE event interval is missing or inverted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Attendance event Start and Stop are required, and Start may not be after Stop.&lt;br /&gt;
One source row has no Start and other rows have Start after Stop.  Production&lt;br /&gt;
attendance events cannot represent an inferred cross-midnight interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE attendance.a_time_start IS NULL&lt;br /&gt;
      OR attendance.a_time_end IS NULL&lt;br /&gt;
     OR attendance.a_time_start &amp;amp;gt; attendance.a_time_end&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned 28 rows: one missing Start, no missing Stop,&lt;br /&gt;
and 27 inverted intervals.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Do not swap times, infer dates, derive Start from duration, or clamp a&lt;br /&gt;
value.  Include the query beside the Problem #175 predicate.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_175&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 28 rows are audit&lt;br /&gt;
evidence; sanity proves every refreshed eligible interval has both endpoints&lt;br /&gt;
and is ordered.  Remove only when all source intervals satisfy that rule.&lt;br /&gt;
&lt;br /&gt;
== * (#176) ATTENDANCE duration disagrees with the event interval ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
For otherwise ordered intervals, some stored durations differ from the exact&lt;br /&gt;
minutes between Start and Stop.  Production has no independent attendance&lt;br /&gt;
duration destination, so loading only the times would discard a conflicting&lt;br /&gt;
source assertion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*,&lt;br /&gt;
       EXTRACT(EPOCH FROM&lt;br /&gt;
               (attendance.a_time_end - attendance.a_time_start)) / 60&lt;br /&gt;
         AS computed_minutes&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE attendance.a_time_start IS NOT NULL&lt;br /&gt;
    AND attendance.a_time_start &amp;amp;lt;= attendance.a_time_end&lt;br /&gt;
    AND attendance.a_duration_of_obs IS DISTINCT FROM&lt;br /&gt;
        EXTRACT(EPOCH FROM&lt;br /&gt;
                (attendance.a_time_end - attendance.a_time_start)) / 60&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned 15 rows, separate from Problem #175.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Do not alter event times or discard the disagreement as warning-only.&lt;br /&gt;
Include the query beside the Problem #176 predicate and prove exact agreement&lt;br /&gt;
for every eligible row.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_176&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 15 rows are audit&lt;br /&gt;
evidence; sanity proves every refreshed eligible duration agrees exactly with&lt;br /&gt;
elapsed time.  Remove only when every otherwise valid source row agrees.&lt;br /&gt;
&lt;br /&gt;
== * (#177) ATTENDANCE intervals overlap for one animal and date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production warns when one animal is recorded at the feeding station in&lt;br /&gt;
overlapping intervals.  Exact duplicates and independently invalid rows must&lt;br /&gt;
be removed first so this issue identifies only otherwise eligible overlap&lt;br /&gt;
participants.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;This query is intentionally complete.  Keep its independent eligibility&lt;br /&gt;
predicates synchronized with Problems #165 and #168 through #176/#178.&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source AS (&lt;br /&gt;
  SELECT row_number() OVER () AS source_id,&lt;br /&gt;
         attendance.*,&lt;br /&gt;
         COUNT(*) OVER (&lt;br /&gt;
           PARTITION BY attendance.a_date,&lt;br /&gt;
                        attendance.a_cl_community_id,&lt;br /&gt;
                        attendance.a_b_animid,&lt;br /&gt;
                        attendance.a_seq_num,&lt;br /&gt;
                        attendance.a_type_of_cycle,&lt;br /&gt;
                        attendance.a_observer1_id,&lt;br /&gt;
                        attendance.a_observer2_id,&lt;br /&gt;
                        attendance.a_bananas_given,&lt;br /&gt;
                        attendance.a_degree_of_arrival,&lt;br /&gt;
                        attendance.a_degree_of_departure,&lt;br /&gt;
                        attendance.a_time_start,&lt;br /&gt;
                        attendance.a_time_end,&lt;br /&gt;
                        attendance.a_duration_of_obs,&lt;br /&gt;
                        attendance.day,&lt;br /&gt;
                        attendance.mo,&lt;br /&gt;
                        attendance.yr,&lt;br /&gt;
                        attendance.cycle_old&lt;br /&gt;
         ) AS duplicate_count&lt;br /&gt;
    FROM clean.attendance AS attendance&lt;br /&gt;
),&lt;br /&gt;
independent_eligible AS (&lt;br /&gt;
  SELECT source.*&lt;br /&gt;
    FROM source&lt;br /&gt;
      JOIN sokwedb.biography_data AS biography&lt;br /&gt;
        ON biography.animid = source.a_b_animid::text&lt;br /&gt;
    WHERE source.duplicate_count = 1&lt;br /&gt;
      AND source.a_date BETWEEN biography.entrydate AND biography.departdate&lt;br /&gt;
      AND COALESCE(BTRIM(source.a_observer1_id::text), &amp;#039;NONE&amp;#039;)&lt;br /&gt;
            &amp;amp;lt;&amp;amp;gt; COALESCE(BTRIM(source.a_observer2_id::text), &amp;#039;NONE&amp;#039;)&lt;br /&gt;
      AND BTRIM(source.a_type_of_cycle::text)&lt;br /&gt;
            NOT IN (&amp;#039;-&amp;#039;, &amp;#039;0.30&amp;#039;, &amp;#039;0.33&amp;#039;)&lt;br /&gt;
      AND source.cycle_old IS NOT NULL&lt;br /&gt;
      AND source.cycle_old::text = BTRIM(source.cycle_old::text)&lt;br /&gt;
      AND (source.a_bananas_given IS NULL&lt;br /&gt;
           OR source.a_bananas_given = TRUNC(source.a_bananas_given))&lt;br /&gt;
      AND source.a_time_start IS NOT NULL&lt;br /&gt;
      AND source.a_time_end IS NOT NULL&lt;br /&gt;
      AND source.a_time_start &amp;amp;lt;= source.a_time_end&lt;br /&gt;
      AND source.a_duration_of_obs IS NOT DISTINCT FROM&lt;br /&gt;
            EXTRACT(EPOCH FROM&lt;br /&gt;
                    (source.a_time_end - source.a_time_start)) / 60&lt;br /&gt;
      AND source.day IS NOT DISTINCT FROM&lt;br /&gt;
        EXTRACT(DAY FROM source.a_date)::integer&lt;br /&gt;
      AND source.mo IS NOT DISTINCT FROM&lt;br /&gt;
        EXTRACT(MONTH FROM source.a_date)::integer&lt;br /&gt;
      AND source.yr IS NOT DISTINCT FROM&lt;br /&gt;
        EXTRACT(YEAR FROM source.a_date)::integer&lt;br /&gt;
)&lt;br /&gt;
SELECT DISTINCT first_row.*&lt;br /&gt;
  FROM independent_eligible AS first_row&lt;br /&gt;
    JOIN independent_eligible AS second_row&lt;br /&gt;
      ON second_row.source_id &amp;amp;lt;&amp;amp;gt; first_row.source_id&lt;br /&gt;
     AND second_row.a_date = first_row.a_date&lt;br /&gt;
     AND second_row.a_b_animid = first_row.a_b_animid&lt;br /&gt;
     AND second_row.a_time_start &amp;amp;lt;= first_row.a_time_end&lt;br /&gt;
     AND second_row.a_time_end &amp;amp;gt;= first_row.a_time_start&lt;br /&gt;
  ORDER BY first_row.a_date,&lt;br /&gt;
           first_row.a_b_animid,&lt;br /&gt;
           first_row.a_time_start,&lt;br /&gt;
           first_row.a_time_end;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed projection returned 428 otherwise eligible source rows after&lt;br /&gt;
Problem #171 became a mapping rather than an exclusion.&lt;br /&gt;
The unfiltered source had 463 overlap pairs, a count that includes other issue&lt;br /&gt;
classes and is not the exclusion baseline.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exclude every row participating in&lt;br /&gt;
an otherwise eligible overlap.  Do not select one interval from a pair.  The&lt;br /&gt;
shared projection must define an execution-local source-row identity, include&lt;br /&gt;
this complete query in code, and prove zero overlaps remain before Problem&lt;br /&gt;
#166 sequence closure.&lt;br /&gt;
&lt;br /&gt;
Implemented after independent exclusions as &amp;lt;code&amp;gt;problem_177&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 428 rows are audit&lt;br /&gt;
evidence; sanity proves that no refreshed eligible intervals overlap.  Remove&lt;br /&gt;
only when no otherwise eligible intervals overlap.&lt;br /&gt;
&lt;br /&gt;
== * (#178) ATTENDANCE date components disagree with A_date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The redundant source Day, Mo, and Yr values should describe&lt;br /&gt;
&amp;lt;code&amp;gt;A_date&amp;lt;/code&amp;gt;.  Some day or month values disagree.  Discarding the&lt;br /&gt;
redundant values would hide a source conflict.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
    WHERE attendance.day IS DISTINCT FROM&lt;br /&gt;
        EXTRACT(DAY FROM attendance.a_date)::integer&lt;br /&gt;
      OR attendance.mo IS DISTINCT FROM&lt;br /&gt;
        EXTRACT(MONTH FROM attendance.a_date)::integer&lt;br /&gt;
      OR attendance.yr IS DISTINCT FROM&lt;br /&gt;
        EXTRACT(YEAR FROM attendance.a_date)::integer&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned 79 distinct rows.  Day disagreed in 46 rows,&lt;br /&gt;
month in 33, and year in none.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Do not silently prefer one component representation.  Include the query&lt;br /&gt;
beside the Problem #178 predicate and prove all eligible redundant components&lt;br /&gt;
agree with &amp;lt;code&amp;gt;A_date&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_178&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 79 rows are audit&lt;br /&gt;
evidence; sanity proves every refreshed eligible redundant date component&lt;br /&gt;
agrees with &amp;lt;code&amp;gt;A_date&amp;lt;/code&amp;gt;.  Remove only when every source row agrees.&lt;br /&gt;
&lt;br /&gt;
Problems #165, #166, #168 through #170, and #172 through #178 overlap and must&lt;br /&gt;
not be summed.  Problem #171 is profiled but mapped, not excluded.  In the&lt;br /&gt;
approved dependency order -- #165, independent #168-#170/#172-#176/#178,&lt;br /&gt;
#177, then #166 -- the refreshed union excluded 13,714 rows and left 120,385&lt;br /&gt;
eligible rows.  Refresh each issue, the overlap set, the final sequence closure,&lt;br /&gt;
and the union before implementation.&lt;br /&gt;
&lt;br /&gt;
On 2026-09-15 the investigator approved replacing complete-row fingerprints&lt;br /&gt;
and frozen issue-membership hashes with semantic validation.  The documented&lt;br /&gt;
predicates and production contracts now determine acceptance.  The conversion&lt;br /&gt;
uses an execution-local &amp;lt;code&amp;gt;source_id&amp;lt;/code&amp;gt; to preserve row multiplicity and&lt;br /&gt;
prove relational parity; it is not a durable source identity.  Historical&lt;br /&gt;
counts remain dated audit evidence and must be refreshed with&lt;br /&gt;
&amp;lt;code&amp;gt;attendance_profile&amp;lt;/code&amp;gt;, including rebuilding the handoff Encounter Log&lt;br /&gt;
when testing a refreshed source.&lt;br /&gt;
&lt;br /&gt;
== (#179) ATTENDANCE observers require standardized PEOPLE matching ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Exact matching would create separate &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; rows for case and&lt;br /&gt;
edge-whitespace variants of the same observer identifier.  This contradicts&lt;br /&gt;
the case-equivalent unique index and the established conversion convention.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Run against the refreshed &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;codes&amp;lt;/code&amp;gt; snapshot after&lt;br /&gt;
food support loading:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH raw_observers AS (&lt;br /&gt;
  SELECT a_observer1_id::text AS raw_person&lt;br /&gt;
    FROM clean.attendance&lt;br /&gt;
    WHERE a_observer1_id IS NOT NULL&lt;br /&gt;
  UNION&lt;br /&gt;
  SELECT a_observer2_id::text&lt;br /&gt;
    FROM clean.attendance&lt;br /&gt;
    WHERE a_observer2_id IS NOT NULL&lt;br /&gt;
), standardized AS (&lt;br /&gt;
  SELECT raw_person,&lt;br /&gt;
         LOWER(NORMALIZE(BTRIM(raw_person))) AS standard_person&lt;br /&gt;
    FROM raw_observers&lt;br /&gt;
)&lt;br /&gt;
SELECT standardized.standard_person,&lt;br /&gt;
       ARRAY_AGG(standardized.raw_person&lt;br /&gt;
                 ORDER BY standardized.raw_person) AS raw_spellings,&lt;br /&gt;
       ARRAY_AGG(people.person ORDER BY people.person)&lt;br /&gt;
         FILTER (WHERE people.person IS NOT NULL) AS existing_codes&lt;br /&gt;
  FROM standardized&lt;br /&gt;
  LEFT JOIN codes.people&lt;br /&gt;
    ON LOWER(NORMALIZE(BTRIM(people.person))) =&lt;br /&gt;
         standardized.standard_person&lt;br /&gt;
  GROUP BY standardized.standard_person&lt;br /&gt;
  ORDER BY standardized.standard_person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The complete refreshed profile contained 932 raw spellings in 718 standardized&lt;br /&gt;
classes.  Sixteen spellings had edge whitespace and 305 raw spellings mapped&lt;br /&gt;
to a different canonical code after trimming and case-equivalent resolution.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: retain the case-equivalent&lt;br /&gt;
&amp;lt;code&amp;gt;people_person_uniquenocase&amp;lt;/code&amp;gt; index and the edge-whitespace constraint.&lt;br /&gt;
Map every raw attendance spelling through the standardized key defined in&lt;br /&gt;
Problem #167.  Existing &amp;lt;code&amp;gt;PEOPLE.Person&amp;lt;/code&amp;gt; spellings take precedence;&lt;br /&gt;
each previously unseen standardized class receives one deterministic support&lt;br /&gt;
code.  Sanity must prove every raw spelling maps exactly once and every&lt;br /&gt;
standardized class produces exactly one canonical code.  This issue excludes&lt;br /&gt;
no rows except the newly recognized standardized observer equalities under&lt;br /&gt;
Problem #168.&lt;br /&gt;
&lt;br /&gt;
== (#180) load_finish deactivates ATTENDANCE observer codes ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The existing finalization step deactivates &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; and every&lt;br /&gt;
&amp;lt;code&amp;gt;PEOPLE.Person&amp;lt;/code&amp;gt; containing a slash.  After attendance loading this&lt;br /&gt;
makes three referenced observer codes inactive: &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;N/A&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;SELEMANI/YAHAYA&amp;lt;/code&amp;gt;.  They account for 37,655&lt;br /&gt;
Recorder or Observer2 uses, violating the requirement that attendance observer&lt;br /&gt;
codes remain active.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Run after attendance and &amp;lt;code&amp;gt;load_finish&amp;lt;/code&amp;gt;:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT people.person,&lt;br /&gt;
       COUNT(*) AS observer_uses&lt;br /&gt;
  FROM (&lt;br /&gt;
    SELECT arrivals_a.recorder AS person FROM sokwedb.arrivals_a&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT arrivals_a.observer2 FROM sokwedb.arrivals_a&lt;br /&gt;
  ) AS attendance_observers&lt;br /&gt;
    JOIN codes.people USING (person)&lt;br /&gt;
  WHERE NOT people.active&lt;br /&gt;
  GROUP BY people.person&lt;br /&gt;
  ORDER BY people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed query returned three codes and 37,655 observer uses:&lt;br /&gt;
&amp;lt;code&amp;gt;N/A&amp;lt;/code&amp;gt; 54 times, &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; 37,599 times, and&lt;br /&gt;
&amp;lt;code&amp;gt;SELEMANI/YAHAYA&amp;lt;/code&amp;gt; twice.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14 as part of the active raw-observer&lt;br /&gt;
policy: &amp;lt;code&amp;gt;conversion/load_finish.sql&amp;lt;/code&amp;gt; retains active status for any&lt;br /&gt;
person referenced by &amp;lt;code&amp;gt;ARRIVALS_A.Recorder&amp;lt;/code&amp;gt; or&lt;br /&gt;
&amp;lt;code&amp;gt;ARRIVALS_A.Observer2&amp;lt;/code&amp;gt;.  Problem #39 finalization remains unchanged&lt;br /&gt;
for unreferenced &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;, and slash-containing codes.&lt;br /&gt;
This issue excludes no attendance rows.  Remove the exception only when no&lt;br /&gt;
attendance detail references a code that finalization would deactivate.&lt;br /&gt;
&lt;br /&gt;
== (#181) COLOBUS descriptive text fields contain NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access &amp;lt;code&amp;gt;COLOBUS&amp;lt;/code&amp;gt; source contains NULL in eight descriptive text&lt;br /&gt;
fields whose production destinations are required text columns.  NULL occurs&lt;br /&gt;
in &amp;lt;code&amp;gt;col_hunter_id&amp;lt;/code&amp;gt; 722 times, &amp;lt;code&amp;gt;col_killer_id&amp;lt;/code&amp;gt; 1,438 times,&lt;br /&gt;
&amp;lt;code&amp;gt;col_victim_age&amp;lt;/code&amp;gt; 1,447 times, &amp;lt;code&amp;gt;col_colobus_details&amp;lt;/code&amp;gt; 1,652 times,&lt;br /&gt;
&amp;lt;code&amp;gt;col_veg_details&amp;lt;/code&amp;gt; 2,073 times, &amp;lt;code&amp;gt;col_hunt_failure_details&amp;lt;/code&amp;gt; 1,943 times,&lt;br /&gt;
&amp;lt;code&amp;gt;col_comments&amp;lt;/code&amp;gt; 885 times, and &amp;lt;code&amp;gt;col_hunt_description&amp;lt;/code&amp;gt;&lt;br /&gt;
1,803 times.&lt;br /&gt;
&lt;br /&gt;
The seven colobus-specific values map to &amp;lt;code&amp;gt;COLOBUS.Hunters&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;COLOBUS.Killers&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;COLOBUS.VictimsAges&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;COLOBUS.Details&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;COLOBUS.Vegetation&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;COLOBUS.FailureDetails&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;COLOBUS.Description&amp;lt;/code&amp;gt;.&lt;br /&gt;
Comments map separately to the shared &amp;lt;code&amp;gt;EVENTS.Notes&amp;lt;/code&amp;gt; column.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (WHERE col_hunter_id IS NULL) AS hunters,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_killer_id IS NULL) AS killers,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_victim_age IS NULL) AS victims_ages,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_colobus_details IS NULL) AS troop_details,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_veg_details IS NULL) AS vegetation,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_hunt_failure_details IS NULL)&lt;br /&gt;
         AS failure_details,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_comments IS NULL) AS comments,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_hunt_description IS NULL) AS descriptions&lt;br /&gt;
  FROM clean.colobus;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed query returned, in column order, 722, 1,438, 1,447, 1,652,&lt;br /&gt;
2,073, 1,943, 885, and 1,803 rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-15: retain the existing production&lt;br /&gt;
schema and map each NULL independently to the empty string.  Implement each&lt;br /&gt;
destination as &amp;lt;code&amp;gt;COALESCE(source_value, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt;; preserve every non-NULL&lt;br /&gt;
source value exactly.  The existing required-text and not-only-spaces&lt;br /&gt;
constraints permit the empty string, so this normalization requires no schema&lt;br /&gt;
change.  Do not concatenate comments with the encounter description, and do&lt;br /&gt;
not exclude a source row solely because any of these eight fields is NULL.&lt;br /&gt;
&lt;br /&gt;
Sanity must prove exact field-by-field &amp;lt;code&amp;gt;COALESCE(source_value, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt;&lt;br /&gt;
parity and zero exclusions attributable only to Problem #181.  Remove this&lt;br /&gt;
normalization only when all eight source fields contain no NULL values.&lt;br /&gt;
&lt;br /&gt;
== (#182) MATING_EVENT community codes differ only by case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access &amp;lt;code&amp;gt;MATING_EVENT.M_cl_community_id&amp;lt;/code&amp;gt; column contains the value &amp;lt;code&amp;gt;kk&amp;lt;/code&amp;gt;,&lt;br /&gt;
which differs only by case from the existing production &amp;lt;code&amp;gt;COMM_IDS.CommID&amp;lt;/code&amp;gt; code&lt;br /&gt;
&amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt;.  Production community codes are case-sensitive.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE m_cl_community_id = &amp;#039;kk&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains 181 such rows.  All other mating community&lt;br /&gt;
values are &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;MT&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigators on 2026-09-15: standardize&lt;br /&gt;
&amp;lt;code&amp;gt;m_cl_community_id = &amp;#039;kk&amp;#039;&amp;lt;/code&amp;gt; to the existing &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt; code in the clean schema.  This&lt;br /&gt;
is a case-only spelling normalization and excludes no rows.  Preserve all&lt;br /&gt;
other source community values unchanged, and require the mating sanity check&lt;br /&gt;
to verify that every resulting community exists in &amp;lt;code&amp;gt;COMM_IDS&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#183) MATING_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;EVENTS.Start&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;EVENTS.Stop&amp;lt;/code&amp;gt; record times to minute precision&lt;br /&gt;
and reject nonzero seconds.  The Access &amp;lt;code&amp;gt;MATING_EVENT.M_time&amp;lt;/code&amp;gt; value is initially&lt;br /&gt;
represented as a timestamp, and merely converting it to a PostgreSQL time&lt;br /&gt;
value does not remove seconds.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;MATING_EVENT&amp;quot;&lt;br /&gt;
  WHERE extract(second FROM &amp;quot;M_time&amp;quot;) &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains one row, dated 2008-10-12 with focal GA, male&lt;br /&gt;
WL, female GA, and time 09:35:30.  Problem #4 remains corrected: the refreshed&lt;br /&gt;
source has no non-midnight time component in &amp;lt;code&amp;gt;M_FOL_date&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigators on 2026-09-15: mating event times are recorded&lt;br /&gt;
to minute precision.  During construction of the tidy schema, truncate&lt;br /&gt;
&amp;lt;code&amp;gt;M_time&amp;lt;/code&amp;gt; to the minute before &amp;lt;code&amp;gt;tidy.sql&amp;lt;/code&amp;gt; converts the column to &amp;lt;code&amp;gt;TIME&amp;lt;/code&amp;gt;.  Do not&lt;br /&gt;
round the value.  The mating sanity check must reject the conversion if a&lt;br /&gt;
value containing seconds nevertheless reaches the clean schema.  This issue&lt;br /&gt;
excludes no rows and does not resolve missing or out-of-window mating times.&lt;br /&gt;
&lt;br /&gt;
== * (#184) MATING_EVENT ordinary boolean flags contain question marks ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;MATINGS.Incest&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;Consort&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;Guarding&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;Courting&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;Camp&amp;lt;/code&amp;gt; are&lt;br /&gt;
required Boolean values.  The Access fields use &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt; for true and blank or NULL&lt;br /&gt;
for false, but one &amp;lt;code&amp;gt;M_court_flag&amp;lt;/code&amp;gt; row and two different &amp;lt;code&amp;gt;M_camp_flag&amp;lt;/code&amp;gt; rows&lt;br /&gt;
contain &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt;.  The unknown values cannot be represented by required Booleans.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE BTRIM(m_court_flag) = &amp;#039;?&amp;#039;&lt;br /&gt;
        OR BTRIM(m_camp_flag) = &amp;#039;?&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains three distinct rows.  No row contains &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt; in&lt;br /&gt;
both fields, and none overlaps the Problem #186 exclusion.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigators on 2026-09-15: map &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt; to true and blank or NULL&lt;br /&gt;
to false for all five fields.  Exclude the three rows containing &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt;.  Sanity&lt;br /&gt;
must reject any other domain value.  This issue independently excludes three&lt;br /&gt;
rows and contributes three rows to the exclusion union.&lt;br /&gt;
&lt;br /&gt;
== (#185) MATING_EVENT records two interference facts in one target field ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;M_interference&amp;lt;/code&amp;gt; records a description or chimp ID.&lt;br /&gt;
&amp;lt;code&amp;gt;M_interference_too_late_flag&amp;lt;/code&amp;gt; records the ID of a chimp who interfered too&lt;br /&gt;
late.  Production has only the single required text column&lt;br /&gt;
&amp;lt;code&amp;gt;MATINGS.Interference&amp;lt;/code&amp;gt;, so copying or concatenating the fields without labels&lt;br /&gt;
would lose which fact each value represents.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (&lt;br /&gt;
         WHERE NULLIF(BTRIM(m_interference), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
               AND NULLIF(BTRIM(m_interference_too_late_flag), &amp;#039;&amp;#039;) IS NOT NULL)&lt;br /&gt;
         AS both_fields,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE NULLIF(BTRIM(m_interference), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
               AND NULLIF(BTRIM(m_interference_too_late_flag), &amp;#039;&amp;#039;) IS NULL)&lt;br /&gt;
         AS interference_only,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE NULLIF(BTRIM(m_interference), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
               AND NULLIF(BTRIM(m_interference_too_late_flag), &amp;#039;&amp;#039;) IS NOT NULL)&lt;br /&gt;
         AS too_late_only&lt;br /&gt;
  FROM clean.mating_event;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains 82 rows with both fields, 3,760 with only&lt;br /&gt;
&amp;lt;code&amp;gt;m_interference&amp;lt;/code&amp;gt;, and 135 with only the too-late field.  No source&lt;br /&gt;
&amp;lt;code&amp;gt;m_interference&amp;lt;/code&amp;gt; contains the labels &amp;lt;code&amp;gt;Interference:&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;Too late:&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: trim both fields and use this&lt;br /&gt;
explicit encoding in &amp;lt;code&amp;gt;MATINGS.Interference&amp;lt;/code&amp;gt;:&lt;br /&gt;
&lt;br /&gt;
* neither field: the empty string;&lt;br /&gt;
* &amp;lt;code&amp;gt;m_interference&amp;lt;/code&amp;gt; only: the trimmed source value;&lt;br /&gt;
* too-late only: &amp;lt;code&amp;gt;Too late: &amp;lt;/code&amp;gt; followed by the trimmed source value; and&lt;br /&gt;
* both fields: &amp;lt;code&amp;gt;Interference: &amp;lt;/code&amp;gt; followed by the trimmed interference value,&lt;br /&gt;
  then &amp;lt;code&amp;gt;; Too late: &amp;lt;/code&amp;gt; followed by the trimmed too-late value.&lt;br /&gt;
&lt;br /&gt;
This preserves the two source facts without changing the production schema.&lt;br /&gt;
Problem #186 separately excludes too-late &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt; values before this encoding.&lt;br /&gt;
This issue excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== * (#186) MATING_EVENT too-late interference contains X instead of an ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The investigators identified &amp;lt;code&amp;gt;M_interference_too_late_flag&amp;lt;/code&amp;gt; as the ID of a&lt;br /&gt;
chimp who interfered too late, but 110 source rows contain &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;.  &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt; is not a&lt;br /&gt;
chimp ID, and the source does not identify an individual who can replace it.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE BTRIM(m_interference_too_late_flag) = &amp;#039;X&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 110 rows span 1976 through 2015 and all use source &amp;lt;code&amp;gt;B-REC&amp;lt;/code&amp;gt;.  Eighty-two&lt;br /&gt;
also contain a separate interference description or ID; 28 do not.  In&lt;br /&gt;
contrast, the 107 chimp-like values occur only in 2012 through 2015 and always&lt;br /&gt;
have an empty main interference field.  The evidence is consistent with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;&lt;br /&gt;
being an older marker, but it does not recover the required chimp ID.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: exclude all 110 rows containing&lt;br /&gt;
&amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt; in &amp;lt;code&amp;gt;m_interference_too_late_flag&amp;lt;/code&amp;gt;.  Do not interpret &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt; as an individual&lt;br /&gt;
or silently discard the too-late fact.  This issue independently excludes 110&lt;br /&gt;
rows, overlaps none of the three Problem #184 rows, and increases the&lt;br /&gt;
exclusion union from 3 to 113 rows.&lt;br /&gt;
&lt;br /&gt;
== * (#187) MATING_EVENT fail flags contain missing and unknown values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;MATINGS.Fail&amp;lt;/code&amp;gt; is a required Boolean.  The investigators confirmed&lt;br /&gt;
that &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;Y&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;y&amp;lt;/code&amp;gt; mean the mating was not complete and map to true, while&lt;br /&gt;
&amp;lt;code&amp;gt;N&amp;lt;/code&amp;gt; maps to false.  The Access source also contains blank, NULL, and &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt;&lt;br /&gt;
values, which require an explicit disposition.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(m_fail_flag), &amp;#039;&amp;#039;), &amp;#039;[blank/NULL]&amp;#039;) AS fail_value,&lt;br /&gt;
       COUNT(*)&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  GROUP BY COALESCE(NULLIF(BTRIM(m_fail_flag), &amp;#039;&amp;#039;), &amp;#039;[blank/NULL]&amp;#039;)&lt;br /&gt;
  ORDER BY fail_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains 21,786 blank or NULL values and 552 &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt; values.&lt;br /&gt;
After Problems #184 and #186, 21,676 blank or NULL values and 551 &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt; values&lt;br /&gt;
remain eligible.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: map &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;Y&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;y&amp;lt;/code&amp;gt; to true;&lt;br /&gt;
map &amp;lt;code&amp;gt;N&amp;lt;/code&amp;gt;, blank, and NULL to false; and exclude rows containing &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt;.  Sanity&lt;br /&gt;
must reject any other domain value.  The &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt; condition independently excludes&lt;br /&gt;
552 rows, overlaps one Problem #184 row and no Problem #186 rows, and increases&lt;br /&gt;
the exclusion union from 113 to 664 rows.  It contributes 551 newly excluded&lt;br /&gt;
rows.&lt;br /&gt;
&lt;br /&gt;
== * (#188) MATING_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the separately governed &amp;lt;code&amp;gt;sdb_no_time&amp;lt;/code&amp;gt; sentinel, production&lt;br /&gt;
&amp;lt;code&amp;gt;EVENTS.Start&amp;lt;/code&amp;gt; cannot be before 04:00 and &amp;lt;code&amp;gt;EVENTS.Stop&amp;lt;/code&amp;gt; cannot be after 20:00.&lt;br /&gt;
Some non-NULL Access mating times fall outside that interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE m_time IS NOT NULL&lt;br /&gt;
        AND (m_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME OR m_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
  ORDER BY m_fol_date, m_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains 25 rows: 24 between 01:00 and 03:48 and one at&lt;br /&gt;
21:38.  None overlaps Problems #184, #186, or #187.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: exclude the 25 rows with non-NULL&lt;br /&gt;
times outside 04:00 through 20:00.  Do not clamp, wrap, or reinterpret their&lt;br /&gt;
times.  This issue independently and newly excludes 25 rows, increasing the&lt;br /&gt;
exclusion union from 664 to 689 rows.&lt;br /&gt;
&lt;br /&gt;
== (#189) MATING_EVENT rows may have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access source contains 34 mating rows with NULL &amp;lt;code&amp;gt;m_time&amp;lt;/code&amp;gt;.  Production&lt;br /&gt;
&amp;lt;code&amp;gt;EVENTS.Start&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;EVENTS.Stop&amp;lt;/code&amp;gt; are required.  Before this decision, only&lt;br /&gt;
aggression, B-record-note, and pantgrunt events could use the established&lt;br /&gt;
&amp;lt;code&amp;gt;sdb_no_time&amp;lt;/code&amp;gt; value to distinguish an unrecorded time from an observed time.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE m_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains 34 rows.  One overlaps the prior exclusion&lt;br /&gt;
union, so excluding missing times would discard 33 additional otherwise&lt;br /&gt;
eligible mating records.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: permit mating events to use the&lt;br /&gt;
existing &amp;lt;code&amp;gt;sdb_no_time&amp;lt;/code&amp;gt; sentinel.  Add &amp;lt;code&amp;gt;sdb_mating_event&amp;lt;/code&amp;gt; to the two owning&lt;br /&gt;
&amp;lt;code&amp;gt;EVENTS&amp;lt;/code&amp;gt; constraint allowlists and document the expanded behavior set.  At the&lt;br /&gt;
production load boundary, map only SQL NULL &amp;lt;code&amp;gt;m_time&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;sdb_no_time&amp;lt;/code&amp;gt;&lt;br /&gt;
for both &amp;lt;code&amp;gt;EVENTS.Start&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;EVENTS.Stop&amp;lt;/code&amp;gt;; preserve every non-NULL source time.&lt;br /&gt;
The existing point-event and paired-sentinel constraints remain in force.&lt;br /&gt;
This issue excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#190) MATING_EVENT swelling values may be missing ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;MATINGS.Swelling&amp;lt;/code&amp;gt; is required, but the Access source contains 572&lt;br /&gt;
NULL values and 16 &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt; values.  The source context for the question marks&lt;br /&gt;
describes unavailable information, including &amp;quot;no estrous state noted&amp;quot; and&lt;br /&gt;
&amp;quot;swelling not recorded&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE m_fem_swelling IS NULL&lt;br /&gt;
        OR BTRIM(m_fem_swelling) = &amp;#039;?&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After prior exclusions, 568 NULL values and all 16 question marks remain&lt;br /&gt;
eligible.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: map both NULL and &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt; to the&lt;br /&gt;
existing &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; code &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, documented as &amp;quot;Missing data&amp;quot;.  &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; has&lt;br /&gt;
NULL &amp;lt;code&amp;gt;AsNum&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;SSRank&amp;lt;/code&amp;gt;, so these records do not contribute a measured value&lt;br /&gt;
to daily swelling minimum/maximum calculations.  The original missing-value&lt;br /&gt;
encoding remains available in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt;.  This issue excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#191) MATING_EVENT records one-third swelling ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Nine otherwise eligible Access mating rows contain the swelling value &amp;lt;code&amp;gt;0.33&amp;lt;/code&amp;gt;,&lt;br /&gt;
which is absent from &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt;.  Inserting the value into &amp;lt;code&amp;gt;MATINGS&amp;lt;/code&amp;gt;&lt;br /&gt;
reproduces a &amp;lt;code&amp;gt;matings_swelling_fkey&amp;lt;/code&amp;gt; violation.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE BTRIM(m_fem_swelling) = &amp;#039;0.33&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: add &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; code &amp;lt;code&amp;gt;0.33&amp;lt;/code&amp;gt;&lt;br /&gt;
with &amp;lt;code&amp;gt;AsNum&amp;lt;/code&amp;gt; 0.33, &amp;lt;code&amp;gt;SSRank&amp;lt;/code&amp;gt; 3, and description &amp;quot;1/3 swollen&amp;quot;.  Shift the&lt;br /&gt;
existing &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; support rows to ranks 4 through 7 so the&lt;br /&gt;
unique integer rank continues to order measured swelling values correctly.&lt;br /&gt;
Map source &amp;lt;code&amp;gt;0.33&amp;lt;/code&amp;gt; to the new exact code.  This issue excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#192) MATING_EVENT extractors require PEOPLE support ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;MATINGS.ExtractedBy&amp;lt;/code&amp;gt; is required and must reference an active&lt;br /&gt;
&amp;lt;code&amp;gt;PEOPLE.Person&amp;lt;/code&amp;gt;.  The source contains 20,747 NULL extractor values, one&lt;br /&gt;
case-only spelling difference, and three populated names absent from the&lt;br /&gt;
conversion &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT m_extracted_by, COUNT(*)&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE NULLIF(BTRIM(m_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
        OR NOT EXISTS (&lt;br /&gt;
             SELECT 1&lt;br /&gt;
               FROM clean.people&lt;br /&gt;
               WHERE LOWER(NORMALIZE(BTRIM(people.person))) =&lt;br /&gt;
                     LOWER(NORMALIZE(BTRIM(mating_event.m_extracted_by))))&lt;br /&gt;
  GROUP BY m_extracted_by&lt;br /&gt;
  ORDER BY m_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;Karen McLellan&amp;lt;/code&amp;gt; matches the existing stored spelling &amp;lt;code&amp;gt;KAREN MCLELLAN&amp;lt;/code&amp;gt;.&lt;br /&gt;
&amp;lt;code&amp;gt;Deus Mjungu&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FELDBLUM&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;Joclyn Antonio&amp;lt;/code&amp;gt; have no standardized match.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: map a NULL, empty, or&lt;br /&gt;
whitespace-only extractor to &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;; trim populated values; reuse the exact&lt;br /&gt;
stored &amp;lt;code&amp;gt;PEOPLE.Person&amp;lt;/code&amp;gt; spelling for standardized case-equivalent matches; and&lt;br /&gt;
add &amp;lt;code&amp;gt;Deus Mjungu&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FELDBLUM&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;Joclyn Antonio&amp;lt;/code&amp;gt; as active support people.&lt;br /&gt;
Use &amp;lt;code&amp;gt;LOWER(NORMALIZE(BTRIM(value)))&amp;lt;/code&amp;gt; as the deterministic matching key.  This&lt;br /&gt;
issue excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#193) MATING_EVENT comments require a production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access source stores mating comments separately from the full mating&lt;br /&gt;
description.  Production has one shared event-notes field, &amp;lt;code&amp;gt;EVENTS.Notes&amp;lt;/code&amp;gt;,&lt;br /&gt;
and silently concatenating the two source fields would erase that distinction.&lt;br /&gt;
The source contains 14,294 NULL comments and 15,134 comments that are NULL or&lt;br /&gt;
empty after trimming.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (WHERE m_comments IS NULL) AS null_comments,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE NULLIF(BTRIM(m_comments), &amp;#039;&amp;#039;) IS NULL) AS blank_comments&lt;br /&gt;
  FROM clean.mating_event;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: map &amp;lt;code&amp;gt;m_comments&amp;lt;/code&amp;gt; alone to&lt;br /&gt;
&amp;lt;code&amp;gt;EVENTS.Notes&amp;lt;/code&amp;gt;.  Map SQL NULL to the empty string with&lt;br /&gt;
&amp;lt;code&amp;gt;COALESCE(m_comments, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt; and preserve every non-NULL source value exactly.&lt;br /&gt;
Do not concatenate &amp;lt;code&amp;gt;m_full_description&amp;lt;/code&amp;gt; into event notes.  This issue excludes&lt;br /&gt;
no rows.&lt;br /&gt;
&lt;br /&gt;
== (#194) MATING_EVENT full descriptions require a mating-specific destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access source stores a full mating description in &amp;lt;code&amp;gt;m_full_description&amp;lt;/code&amp;gt;.&lt;br /&gt;
&amp;lt;code&amp;gt;BRECORD_NOTES&amp;lt;/code&amp;gt; cannot preserve it because that detail table can attach only&lt;br /&gt;
to a B-record event, cannot share the mating EID, and requires unrelated&lt;br /&gt;
B-record text fields.  The source contains 18,212 descriptions that are NULL&lt;br /&gt;
or empty after trimming.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (WHERE m_full_description IS NULL)&lt;br /&gt;
         AS null_descriptions,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE NULLIF(BTRIM(m_full_description), &amp;#039;&amp;#039;) IS NULL)&lt;br /&gt;
         AS blank_descriptions&lt;br /&gt;
  FROM clean.mating_event;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: add the required text column&lt;br /&gt;
&amp;lt;code&amp;gt;MATINGS.Description&amp;lt;/code&amp;gt; and map &amp;lt;code&amp;gt;m_full_description&amp;lt;/code&amp;gt; to it.  Map SQL NULL to the&lt;br /&gt;
empty string with &amp;lt;code&amp;gt;COALESCE(m_full_description, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt; and preserve every&lt;br /&gt;
non-NULL source value exactly.  The production constraint permits the empty&lt;br /&gt;
string but rejects nonempty whitespace-only values.&lt;br /&gt;
&lt;br /&gt;
The permanent schema and documentation change was implemented as append-only&lt;br /&gt;
master commit &amp;lt;code&amp;gt;a628fdd&amp;lt;/code&amp;gt; (&amp;lt;code&amp;gt;feat(matings): preserve mating descriptions&amp;lt;/code&amp;gt;) and&lt;br /&gt;
fast-forwarded to local, origin, private, and blessed master.  This issue&lt;br /&gt;
excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#195) MATING_EVENT sources require MATING_SOURCES support ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;MATINGS.Source&amp;lt;/code&amp;gt; is required and references&lt;br /&gt;
&amp;lt;code&amp;gt;MATING_SOURCES.Source&amp;lt;/code&amp;gt;, but the source code table is empty.  The Access source&lt;br /&gt;
contains mixed-case source labels while the production support key requires&lt;br /&gt;
uppercase values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT m_source, COUNT(*) AS row_count,&lt;br /&gt;
       MIN(m_fol_date) AS first_date,&lt;br /&gt;
       MAX(m_fol_date) AS last_date&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  GROUP BY m_source&lt;br /&gt;
  ORDER BY m_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains &amp;lt;code&amp;gt;B-REC&amp;lt;/code&amp;gt; 26,066 times over 1976 through 2018,&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt; 393 times over 2000 through 2007, and &amp;lt;code&amp;gt;Bobu&amp;lt;/code&amp;gt; 24 times in 2000.  The&lt;br /&gt;
repository provides no more authoritative definition of &amp;lt;code&amp;gt;Bobu&amp;lt;/code&amp;gt; than its&lt;br /&gt;
source label.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: add literal support codes &amp;lt;code&amp;gt;B-REC&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;TIKI&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;BOBU&amp;lt;/code&amp;gt;, described respectively as &amp;quot;Mating record sourced from&lt;br /&gt;
B-REC&amp;quot;, &amp;quot;Mating record sourced from Tiki&amp;quot;, and &amp;quot;Mating record sourced from&lt;br /&gt;
Bobu&amp;quot;.  At the production boundary map source &amp;lt;code&amp;gt;B-REC&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;B-REC&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;TIKI&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;Bobu&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;BOBU&amp;lt;/code&amp;gt;.  Do not claim semantics beyond the source labels.&lt;br /&gt;
Sanity must reject any refreshed source value outside this approved domain and&lt;br /&gt;
verify that every transformed code exists exactly once.  This issue excludes&lt;br /&gt;
no rows.&lt;br /&gt;
&lt;br /&gt;
== * (#196) MATING_EVENT participants are absent from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;ROLES.Participant&amp;lt;/code&amp;gt; is required and references&lt;br /&gt;
&amp;lt;code&amp;gt;BIOGRAPHY_DATA.AnimID&amp;lt;/code&amp;gt;.  The first exclusion-free loader run failed while&lt;br /&gt;
inserting female participant &amp;lt;code&amp;gt;STR&amp;lt;/code&amp;gt; for conversion identity 117.  The source&lt;br /&gt;
also contains other participant values that are not production biography&lt;br /&gt;
identifiers, including category labels, unknown markers, blanks, and&lt;br /&gt;
unresolved spellings.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH eligible AS (&lt;br /&gt;
  SELECT *&lt;br /&gt;
    FROM clean.mating_event&lt;br /&gt;
    WHERE COALESCE(BTRIM(m_court_flag), &amp;#039;&amp;#039;) &amp;lt;&amp;gt; &amp;#039;?&amp;#039;&lt;br /&gt;
          AND COALESCE(BTRIM(m_camp_flag), &amp;#039;&amp;#039;) &amp;lt;&amp;gt; &amp;#039;?&amp;#039;&lt;br /&gt;
          AND COALESCE(BTRIM(m_interference_too_late_flag), &amp;#039;&amp;#039;) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
          AND COALESCE(BTRIM(m_fail_flag), &amp;#039;&amp;#039;) &amp;lt;&amp;gt; &amp;#039;?&amp;#039;&lt;br /&gt;
          AND (m_time IS NULL&lt;br /&gt;
               OR m_time BETWEEN &amp;#039;04:00&amp;#039;::TIME AND &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
)&lt;br /&gt;
SELECT DISTINCT eligible.m_conversion_id&lt;br /&gt;
  FROM eligible&lt;br /&gt;
    CROSS JOIN LATERAL (&lt;br /&gt;
      VALUES (eligible.m_male), (eligible.m_female)&lt;br /&gt;
    ) AS participant(animid)&lt;br /&gt;
  WHERE NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
      FROM sokwedb.biography_data&lt;br /&gt;
      WHERE biography_data.animid = BTRIM(participant.animid))&lt;br /&gt;
  ORDER BY eligible.m_conversion_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed post-follow database contains 574 missing participant&lt;br /&gt;
occurrences in 566 otherwise eligible source rows, representing 34 distinct&lt;br /&gt;
trimmed values.  The first failure is identity 117, dated 2000-10-14, with&lt;br /&gt;
focal &amp;lt;code&amp;gt;BE&amp;lt;/code&amp;gt;, male &amp;lt;code&amp;gt;SL&amp;lt;/code&amp;gt;, and female &amp;lt;code&amp;gt;STR&amp;lt;/code&amp;gt;.  The failed first chunk rolled back&lt;br /&gt;
and left zero &amp;lt;code&amp;gt;MATE&amp;lt;/code&amp;gt; events.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: exclude every source row for which&lt;br /&gt;
either trimmed participant is absent from &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt;.  Do not create&lt;br /&gt;
biography rows or substitute participant identities without source evidence.&lt;br /&gt;
Sanity must verify that this condition selects exactly 566 distinct eligible&lt;br /&gt;
source identities in the retained snapshot.  The rows overlap none of the&lt;br /&gt;
prior 689-row exclusion union and increase it to 1,255 rows, leaving 25,228&lt;br /&gt;
rows eligible before later failures.&lt;br /&gt;
&lt;br /&gt;
== * (#197) MATING_EVENT participants fall outside biography study dates ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After excluding Problem #196 rows, the loader failed when actor &amp;lt;code&amp;gt;FAM&amp;lt;/code&amp;gt;&lt;br /&gt;
participated on 2014-09-10, after the individual&amp;#039;s 2012-04-26 departure date.&lt;br /&gt;
Production requires every role participant to be biography-valid on the event&lt;br /&gt;
date.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT DISTINCT mating_event.m_conversion_id&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
    CROSS JOIN LATERAL (&lt;br /&gt;
      VALUES (mating_event.m_male), (mating_event.m_female)&lt;br /&gt;
    ) AS participant(animid)&lt;br /&gt;
    JOIN sokwedb.biography_data&lt;br /&gt;
      ON biography_data.animid = BTRIM(participant.animid)&lt;br /&gt;
  WHERE mating_event.m_fol_date &amp;lt; biography_data.entrydate&lt;br /&gt;
        OR mating_event.m_fol_date &amp;gt; biography_data.departdate&lt;br /&gt;
  ORDER BY mating_event.m_conversion_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After applying Problems #184, #186, #187, #188, and #196, 28 participant&lt;br /&gt;
occurrences in 28 source rows fall outside study dates.  Three are male actors&lt;br /&gt;
after departure, 13 are female actees after departure, and 12 are female&lt;br /&gt;
actees before entry.  The affected participants are &amp;lt;code&amp;gt;AT&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;CF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FAM&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OR&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;SF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;SH&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;SIF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;SP&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;WN&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YD&amp;lt;/code&amp;gt;.  The failed first chunk rolled back&lt;br /&gt;
and left zero &amp;lt;code&amp;gt;MATE&amp;lt;/code&amp;gt; events.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: exclude all 28 rows.  Do not alter&lt;br /&gt;
source dates or biography study spans.  Sanity must verify the exact 28-row&lt;br /&gt;
condition after the prior exclusions.  This increases the exclusion union&lt;br /&gt;
from 1,255 to 1,283 rows and leaves 25,200 rows eligible before later&lt;br /&gt;
failures.&lt;br /&gt;
&lt;br /&gt;
== * (#198) MATING_EVENT focal values are absent from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After Problems #196 and #197, the loader failed while creating an Other watch&lt;br /&gt;
for focal &amp;lt;code&amp;gt;LUT&amp;lt;/code&amp;gt;, which is absent from &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt;.  Production requires&lt;br /&gt;
every &amp;lt;code&amp;gt;WATCHES.AnimID&amp;lt;/code&amp;gt;, including mating-only Other watches, to be a biography&lt;br /&gt;
identifier or the established no-focal sentinel.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(m_fol_b_focal_animid), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS focal,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
      FROM sokwedb.biography_data&lt;br /&gt;
      WHERE biography_data.animid = COALESCE(&lt;br /&gt;
              NULLIF(BTRIM(mating_event.m_fol_b_focal_animid), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;))&lt;br /&gt;
  GROUP BY COALESCE(NULLIF(BTRIM(m_fol_b_focal_animid), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;)&lt;br /&gt;
  ORDER BY focal;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After the prior exclusions, 311 otherwise eligible rows use 51 unsupported&lt;br /&gt;
focal values.  Three source rows have unique case-insensitive biography&lt;br /&gt;
matches: &amp;lt;code&amp;gt;Gb&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;GB&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;lam&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;LAM&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;sif&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;SIF&amp;lt;/code&amp;gt;.  The remaining 308&lt;br /&gt;
rows use 48 unsupported values, many representing groups, multiple animals,&lt;br /&gt;
or missing-target markers.  The failed first chunk rolled back and left zero&lt;br /&gt;
&amp;lt;code&amp;gt;MATE&amp;lt;/code&amp;gt; events.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: normalize only the three proven&lt;br /&gt;
case variants in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt;, then exclude the 308 rows whose normalized focal is&lt;br /&gt;
still absent from &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt;.  Do not infer individual focal identities&lt;br /&gt;
from group or multi-animal values.  Sanity must verify the exact 308-row&lt;br /&gt;
residual after prior exclusions.  This increases the exclusion union from&lt;br /&gt;
1,283 to 1,591 rows and leaves 24,892 rows eligible before later failures.&lt;br /&gt;
&lt;br /&gt;
== * (#199) MATING_EVENT focal dates are outside BIOGRAPHY_DATA study dates ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After Problem #198, the loader failed while creating an Other watch for focal&lt;br /&gt;
&amp;lt;code&amp;gt;POR&amp;lt;/code&amp;gt; on 2016-09-22.  &amp;lt;code&amp;gt;POR&amp;lt;/code&amp;gt; departed on 2014-03-13, and production requires a&lt;br /&gt;
watch date to fall within the focal animal&amp;#039;s biography study dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT mating_event.m_conversion_id,&lt;br /&gt;
       mating_event.m_fol_date,&lt;br /&gt;
       focal_biography.animid,&lt;br /&gt;
       focal_biography.entrydate,&lt;br /&gt;
       focal_biography.departdate&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
    JOIN sokwedb.biography_data AS focal_biography&lt;br /&gt;
      ON focal_biography.animid = COALESCE(&lt;br /&gt;
           NULLIF(BTRIM(mating_event.m_fol_b_focal_animid), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;)&lt;br /&gt;
  WHERE mating_event.m_fol_date &amp;lt; focal_biography.entrydate&lt;br /&gt;
        OR mating_event.m_fol_date &amp;gt; focal_biography.departdate&lt;br /&gt;
  ORDER BY mating_event.m_conversion_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After the prior exclusions, 43 otherwise eligible rows have focal dates&lt;br /&gt;
outside the focal&amp;#039;s study dates.  Forty-one rows use &amp;lt;code&amp;gt;POR&amp;lt;/code&amp;gt; after departure;&lt;br /&gt;
two use &amp;lt;code&amp;gt;SN&amp;lt;/code&amp;gt; before entry.  The failure occurred after 16,000 rows had&lt;br /&gt;
committed in earlier chunks, while the failing chunk rolled back.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: exclude all 43 rows.  Do not alter&lt;br /&gt;
source dates or biography study spans.  Sanity must verify the exact 43-row&lt;br /&gt;
condition after the prior exclusions.  This increases the exclusion union&lt;br /&gt;
from 1,591 to 1,634 rows and leaves 24,849 rows eligible before later&lt;br /&gt;
failures.  Rebuild the post-follow baseline before retrying so committed MATE&lt;br /&gt;
rows and any mating-created Other watches are removed together.&lt;br /&gt;
&lt;br /&gt;
== (#200) COLOBUS encounters have no recorded Stop ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 204 COLOBUS source rows with a recorded Start but no Stop.  The&lt;br /&gt;
general EVENTS sentinel rules did not permit preserving a known Start with an&lt;br /&gt;
unknown Stop for behavior &amp;lt;code&amp;gt;COL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT col_encounter_date, col_fol_b_focal_chimp_id, col_begin_time&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_end_time IS NULL&lt;br /&gt;
  ORDER BY col_encounter_date, col_fol_b_focal_chimp_id, col_begin_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: retain each recorded Start and map only the&lt;br /&gt;
missing Stop to &amp;lt;code&amp;gt;sdb_no_time&amp;lt;/code&amp;gt;.  Add a narrow, documented EVENTS exception for&lt;br /&gt;
COL with a non-sentinel Start and sentinel Stop.  Do not exclude these rows or&lt;br /&gt;
allow a sentinel COL Start.  Validation must find exactly 204 converted COL&lt;br /&gt;
events with the sentinel Stop.&lt;br /&gt;
&lt;br /&gt;
== * (#201) COLOBUS encounter times violate the approved time range ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Fifteen rows cannot be represented without changing a recorded event time:&lt;br /&gt;
eight have Start after Stop, two have Start after 20:00, and seven have Stop&lt;br /&gt;
after 20:00.  The two late Starts overlap the Start-after-Stop class, so the&lt;br /&gt;
distinct union is 15 rows.  None has a missing Stop.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT col_encounter_date, col_fol_b_focal_chimp_id,&lt;br /&gt;
       col_begin_time, col_end_time&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_begin_time &amp;gt; col_end_time&lt;br /&gt;
        OR col_begin_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME&lt;br /&gt;
        OR col_end_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME&lt;br /&gt;
  ORDER BY col_encounter_date, col_fol_b_focal_chimp_id, col_begin_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: exclude exactly this 15-row union at the loader&lt;br /&gt;
boundary and preserve all other recorded times.  Sanity must fail if the&lt;br /&gt;
predicate no longer identifies exactly 15 rows.  This establishes an&lt;br /&gt;
exclusion union of 15 rows before Problem #210.&lt;br /&gt;
&lt;br /&gt;
== (#202) COLOBUS StartMap is missing from seven rows ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Seven source rows have no &amp;lt;code&amp;gt;COL_begin_time_MAP&amp;lt;/code&amp;gt;.  Of the 2,240 recorded values,&lt;br /&gt;
126 intentionally differ from the standard 15-minute calculation and must not&lt;br /&gt;
be recalculated.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (WHERE col_begin_time_map IS NULL) AS missing,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE col_begin_time_map IS NOT NULL&lt;br /&gt;
               AND col_begin_time_map IS DISTINCT FROM&lt;br /&gt;
                 DATE_BIN(&amp;#039;15 minutes&amp;#039;,&lt;br /&gt;
                          DATE &amp;#039;1960-07-04&amp;#039; + col_begin_time&lt;br /&gt;
                            + INTERVAL &amp;#039;8 minutes&amp;#039;,&lt;br /&gt;
                          TIMESTAMP &amp;#039;1960-07-04 00:00&amp;#039;)::TIME) AS adjusted&lt;br /&gt;
  FROM clean.colobus;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: derive only the seven missing values at the&lt;br /&gt;
production load boundary with the displayed &amp;lt;code&amp;gt;DATE_BIN&amp;lt;/code&amp;gt; expression.  Preserve&lt;br /&gt;
all recorded values exactly.  This excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#203) COLOBUS Hunt, HuntCert, and Kill can be unknown ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The source contains 15 NULL Hunt flags, 16 NULL HuntCert flags, and 221 NULL&lt;br /&gt;
Kill flags, but the corresponding production columns were &amp;lt;code&amp;gt;NOT NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (WHERE col_hunt_flag IS NULL) AS hunt_unknown,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_hunt_cert IS NULL) AS huntcert_unknown,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_kill_flag IS NULL) AS kill_unknown&lt;br /&gt;
  FROM clean.colobus;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: make all three production columns nullable,&lt;br /&gt;
map &amp;lt;code&amp;gt;Y&amp;lt;/code&amp;gt; to true and &amp;lt;code&amp;gt;N&amp;lt;/code&amp;gt; to false, and preserve source NULL as SQL NULL.  This&lt;br /&gt;
excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#204) COLOBUS FocalHuntCert can be unknown ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 744 NULL focal-hunt-certainty values.  Documentation described an&lt;br /&gt;
unknown state, but the production column was &amp;lt;code&amp;gt;NOT NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*)&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_focal_hunt_cert IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: resolve the schema/documentation conflict in&lt;br /&gt;
favor of the documented unknown state.  Make &amp;lt;code&amp;gt;FocalHuntCert&amp;lt;/code&amp;gt; nullable, map&lt;br /&gt;
&amp;lt;code&amp;gt;Y&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;N&amp;lt;/code&amp;gt; to booleans, and preserve all source NULLs.  This excludes no&lt;br /&gt;
rows.&lt;br /&gt;
&lt;br /&gt;
== (#205) COLOBUS NumKills is absent for several distinct states ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 389 rows without a total kill count: seven have &amp;lt;code&amp;gt;Kill = &amp;#039;Y&amp;#039;&amp;lt;/code&amp;gt;, 163&lt;br /&gt;
have &amp;lt;code&amp;gt;Kill = &amp;#039;N&amp;#039;&amp;lt;/code&amp;gt;, and 219 also have an unknown Kill flag.  Seven meaningful&lt;br /&gt;
recorded male totals exceed the recorded overall total and must not be&lt;br /&gt;
rewritten.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT col_kill_flag, COUNT(*)&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_num_kills_total IS NULL&lt;br /&gt;
  GROUP BY col_kill_flag&lt;br /&gt;
  ORDER BY col_kill_flag NULLS LAST;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: make &amp;lt;code&amp;gt;NumKills&amp;lt;/code&amp;gt; nullable.  Preserve every&lt;br /&gt;
recorded total; for a missing total map &amp;lt;code&amp;gt;Kill = &amp;#039;Y&amp;#039;&amp;lt;/code&amp;gt; to 1, &amp;lt;code&amp;gt;Kill = &amp;#039;N&amp;#039;&amp;lt;/code&amp;gt; to 0,&lt;br /&gt;
and unknown Kill to SQL NULL.  Preserve the seven male/total differences and&lt;br /&gt;
verify source-to-target parity.  This excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#206) COLOBUS FemaleNumKills has no supported derivation ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The source does not provide an independently supported total female kill&lt;br /&gt;
count.  Deriving it from other totals would invent data and can be invalid for&lt;br /&gt;
the seven rows where male kills exceed the overall total.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*)&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_num_kills_total IS NOT NULL&lt;br /&gt;
        AND col_num_kills_by_males &amp;gt; col_num_kills_total;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The query returns seven rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: make &amp;lt;code&amp;gt;FemaleNumKills&amp;lt;/code&amp;gt; nullable and always load&lt;br /&gt;
SQL NULL.  Do not derive it by subtraction.  This excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#207) COLOBUS focal kill counts contain fractions ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Three focal-male and three focal-female source counts have fractional values.&lt;br /&gt;
The production &amp;lt;code&amp;gt;INTEGER&amp;lt;/code&amp;gt; columns cannot preserve them.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (&lt;br /&gt;
         WHERE col_num_kills_by_focal_male &amp;lt;&amp;gt;&lt;br /&gt;
               TRUNC(col_num_kills_by_focal_male)) AS focal_male,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE col_num_kills_by_focal_female &amp;lt;&amp;gt;&lt;br /&gt;
               TRUNC(col_num_kills_by_focal_female)) AS focal_female&lt;br /&gt;
  FROM clean.colobus;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: change &amp;lt;code&amp;gt;FocalMaleNumKills&amp;lt;/code&amp;gt; and&lt;br /&gt;
&amp;lt;code&amp;gt;FocalFemaleNumKills&amp;lt;/code&amp;gt; to exact &amp;lt;code&amp;gt;NUMERIC&amp;lt;/code&amp;gt; columns and preserve each source&lt;br /&gt;
value.  This excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#208) COLOBUS group-size codes are unsupported and can be absent ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The support table was empty.  Source values use codes &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt;, and 611&lt;br /&gt;
rows have no group-size code.  Missing values must not be changed to code &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt;,&lt;br /&gt;
which has a distinct recorded meaning.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT col_colobus_group_size, COUNT(*)&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  GROUP BY col_colobus_group_size&lt;br /&gt;
  ORDER BY col_colobus_group_size NULLS LAST;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The counts are 18, 3, 11, 1,555, and 49 for codes &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt;, plus 611&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: load the five documented support rows, retain&lt;br /&gt;
the foreign key, make &amp;lt;code&amp;gt;GroupSize&amp;lt;/code&amp;gt; nullable, and preserve source NULL as SQL&lt;br /&gt;
NULL.  The descriptions are &amp;lt;code&amp;gt;group consisted of only 1 col&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;group consisted&lt;br /&gt;
of only 2 col&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;group defined as small by observer&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;group but no size&lt;br /&gt;
estimate given&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;group defined as large&amp;lt;/code&amp;gt;.  This excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#209) COLOBUS focal identifiers require approved corrections ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Five source rows use four noncanonical focal identifiers.  The approved&lt;br /&gt;
corrections are &amp;lt;code&amp;gt;GRP&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;STR&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;FLT&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;FLI&amp;lt;/code&amp;gt;.  &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; does not fit the source column&amp;#039;s &amp;lt;code&amp;gt;VARCHAR(3)&amp;lt;/code&amp;gt; declaration.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT col_fol_b_focal_chimp_id, COUNT(*)&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_fol_b_focal_chimp_id IN (&amp;#039;GRP&amp;#039;, &amp;#039;STR&amp;#039;, &amp;#039;U&amp;#039;, &amp;#039;FLT&amp;#039;)&lt;br /&gt;
  GROUP BY col_fol_b_focal_chimp_id&lt;br /&gt;
  ORDER BY col_fol_b_focal_chimp_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: widen only the clean copy of the focal column to&lt;br /&gt;
&amp;lt;code&amp;gt;TEXT&amp;lt;/code&amp;gt;, then apply the four corrections to the five affected rows.  Preserve&lt;br /&gt;
raw, tidy, and easy source fidelity.  This correction itself excludes no&lt;br /&gt;
rows; the two corrected &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; rows are handled independently by Problem #210.&lt;br /&gt;
&lt;br /&gt;
== * (#210) Corrected COLOBUS unknown focals have no valid watch context ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After Problem #209 and all preceding watch loaders, the corrected &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; rows&lt;br /&gt;
at 1982-10-30 09:10 and 1985-07-21 10:12 have neither a B/Other watch nor a&lt;br /&gt;
covering &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; row.  Inferring a community from unrelated same-day&lt;br /&gt;
observations would create unsupported data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT col_encounter_date, col_fol_b_focal_chimp_id, col_begin_time&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_fol_b_focal_chimp_id = &amp;#039;UNK&amp;#039;&lt;br /&gt;
        AND (col_encounter_date, col_begin_time) IN (&lt;br /&gt;
              (DATE &amp;#039;1982-10-30&amp;#039;, TIME &amp;#039;09:10&amp;#039;),&lt;br /&gt;
              (DATE &amp;#039;1985-07-21&amp;#039;, TIME &amp;#039;10:12&amp;#039;))&lt;br /&gt;
  ORDER BY col_encounter_date, col_begin_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The investigator approved excluding exactly these two rows.  Do not create a&lt;br /&gt;
watch with an inferred community.  They do not overlap Problem #201, raising&lt;br /&gt;
the exact exclusion union from 15 to 17 rows.  Of 2,247 source rows, 2,230 are&lt;br /&gt;
therefore eligible for conversion.&lt;br /&gt;
&lt;br /&gt;
== * (#211) Corrected COLOBUS no-focal row has conflicting community context ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After Problem #209 and all preceding watch loaders, the corrected &amp;lt;code&amp;gt;GRP&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row at 1978-09-10 11:33 has neither a B/Other watch nor a covering&lt;br /&gt;
&amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; row.  Its source comment says &amp;lt;code&amp;gt;ON PATROL, KAHAMA. NO FOCAL&amp;lt;/code&amp;gt;, but&lt;br /&gt;
the unique same-day attendance community is &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt; (Kasekela); &amp;lt;code&amp;gt;HK&amp;lt;/code&amp;gt; is Kahama&lt;br /&gt;
and its support notes say that community ended on 1977-12-31.  Either community&lt;br /&gt;
choice would therefore infer unsupported data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT col_encounter_date, col_fol_b_focal_chimp_id,&lt;br /&gt;
       col_begin_time, col_comments&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_fol_b_focal_chimp_id = &amp;#039;NONE&amp;#039;&lt;br /&gt;
        AND col_encounter_date = DATE &amp;#039;1978-09-10&amp;#039;&lt;br /&gt;
        AND col_begin_time = TIME &amp;#039;11:33&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The investigator approved excluding exactly this row.  Do not create a&lt;br /&gt;
no-focal Other watch with an inferred community.  It does not overlap Problems&lt;br /&gt;
#201 or #210, raising the exact exclusion union from 17 to 18 rows.  Of 2,247&lt;br /&gt;
source rows, 2,229 are therefore eligible for conversion.&lt;br /&gt;
&lt;br /&gt;
== (#212) OTHER_SPECIES requires production species support codes ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;codes.species&amp;lt;/code&amp;gt; is empty, while &amp;lt;code&amp;gt;clean.other_species_lookup&amp;lt;/code&amp;gt; contains 15&lt;br /&gt;
unique local-species names and unique English descriptions.  Every&lt;br /&gt;
&amp;lt;code&amp;gt;SPECIES_PRESENT.Species&amp;lt;/code&amp;gt; value must reference a nonempty uppercase&lt;br /&gt;
&amp;lt;code&amp;gt;codes.species.Species&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT UPPER(BTRIM(osl_local_species_name)) AS proposed_species,&lt;br /&gt;
       osl_english_species_name AS proposed_description&lt;br /&gt;
  FROM clean.other_species_lookup&lt;br /&gt;
  ORDER BY proposed_species;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The query returns 15 rows: &amp;lt;code&amp;gt;CHATU&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FUNGO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;KAKAKUONA&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;KENGE&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;KIMA&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;MBOGO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;NGURUWE&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;NKUNGE&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;NYANI&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;NYOKA&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;PONGO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;TUMBILI&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;UNKNOWN&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;UNRECORDED&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;VYONDI&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: populate &amp;lt;code&amp;gt;codes.species&amp;lt;/code&amp;gt; using&lt;br /&gt;
the uppercase, trimmed local lookup name as &amp;lt;code&amp;gt;Species&amp;lt;/code&amp;gt; and the corresponding&lt;br /&gt;
English lookup name as &amp;lt;code&amp;gt;Description&amp;lt;/code&amp;gt;.  Preserve all 15 descriptions exactly,&lt;br /&gt;
including &amp;lt;code&amp;gt;Mongoose? Civet? TBD&amp;lt;/code&amp;gt;.  Sanity must verify the exact 15 pairs and&lt;br /&gt;
their uniqueness before loading.  This excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#213) OTHER_SPECIES species labels require canonical standardization ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The source stores 30 distinct spellings of its classified local-species field.&lt;br /&gt;
Trimmed, case-insensitive matching resolves 6,360 of 6,366 rows to the 15-row&lt;br /&gt;
lookup.  One additional row uses &amp;lt;code&amp;gt;Viondi&amp;lt;/code&amp;gt;, a spelling variant of lookup value&lt;br /&gt;
&amp;lt;code&amp;gt;Vyondi&amp;lt;/code&amp;gt;.  The production code must use the uppercase canonical lookup value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT quote_literal(os_osl_local_species_name) AS source_value,&lt;br /&gt;
       COUNT(*) AS rows&lt;br /&gt;
  FROM clean.other_species AS other_species&lt;br /&gt;
  WHERE NOT EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
            FROM clean.other_species_lookup AS lookup&lt;br /&gt;
            WHERE UPPER(BTRIM(lookup.osl_local_species_name)) =&lt;br /&gt;
                  UPPER(BTRIM(other_species.os_osl_local_species_name)))&lt;br /&gt;
  GROUP BY os_osl_local_species_name&lt;br /&gt;
  ORDER BY source_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The query returns &amp;lt;code&amp;gt;Sokwe mgeni&amp;lt;/code&amp;gt; (4), &amp;lt;code&amp;gt;UMBA &amp;lt;/code&amp;gt; (1), and &amp;lt;code&amp;gt;Viondi&amp;lt;/code&amp;gt; (1).&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: standardize classified source&lt;br /&gt;
labels by trimmed, case-insensitive lookup and use the lookup&amp;#039;s uppercase&lt;br /&gt;
local name as the production code.  Correct the single &amp;lt;code&amp;gt;Viondi&amp;lt;/code&amp;gt; value to&lt;br /&gt;
canonical &amp;lt;code&amp;gt;VYONDI&amp;lt;/code&amp;gt;.  Do not use &amp;lt;code&amp;gt;OS_local_species_name_written&amp;lt;/code&amp;gt; to override&lt;br /&gt;
the classified field.  This resolves 6,361 rows and excludes none.  The five&lt;br /&gt;
remaining unlisted rows are handled independently by Problem #214.&lt;br /&gt;
&lt;br /&gt;
== * (#214) OTHER_SPECIES contains unlisted species labels ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Four source rows are classified as &amp;lt;code&amp;gt;Sokwe mgeni&amp;lt;/code&amp;gt;, and one is classified as&lt;br /&gt;
&amp;lt;code&amp;gt;UMBA &amp;lt;/code&amp;gt; with comment &amp;lt;code&amp;gt;&amp;#039;umba&amp;#039; written on sheet&amp;lt;/code&amp;gt;.  Neither trimmed label exists&lt;br /&gt;
in &amp;lt;code&amp;gt;clean.other_species_lookup&amp;lt;/code&amp;gt; or has an approved &amp;lt;code&amp;gt;codes.species&amp;lt;/code&amp;gt; code or&lt;br /&gt;
description.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT os_fol_date, os_fol_b_focal_animid, os_time_begin,&lt;br /&gt;
       os_time_end, os_osl_local_species_name,&lt;br /&gt;
       os_local_species_name_written, os_comments&lt;br /&gt;
  FROM clean.other_species AS other_species&lt;br /&gt;
  WHERE UPPER(BTRIM(os_osl_local_species_name)) IN (&amp;#039;SOKWE MGENI&amp;#039;, &amp;#039;UMBA&amp;#039;)&lt;br /&gt;
  ORDER BY os_fol_date, os_fol_b_focal_animid, os_time_begin;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The five rows overlap neither the Problem #215 focal exclusion nor the Problem&lt;br /&gt;
#216 time exclusion.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: exclude exactly these five rows at&lt;br /&gt;
the loader boundary.  Do not map either label to &amp;lt;code&amp;gt;UNKNOWN&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;UNRECORDED&amp;lt;/code&amp;gt; and&lt;br /&gt;
do not create unsupported species codes.  These rows overlap none of Problems&lt;br /&gt;
#215, #216, or #219 and contribute five rows to the final 13-row exclusion&lt;br /&gt;
union.&lt;br /&gt;
&lt;br /&gt;
== * (#215) OTHER_SPECIES focal identifiers require normalization and one exclusion ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Twelve rows have trailing spaces in &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, one row uses&lt;br /&gt;
lowercase &amp;lt;code&amp;gt;pl&amp;lt;/code&amp;gt;, and one row uses &amp;lt;code&amp;gt;IMB&amp;lt;/code&amp;gt;.  The trimmed values resolve uniquely&lt;br /&gt;
to canonical production IDs for the 12 edge-spaced rows, and &amp;lt;code&amp;gt;pl&amp;lt;/code&amp;gt; resolves&lt;br /&gt;
uniquely to &amp;lt;code&amp;gt;PL&amp;lt;/code&amp;gt;.  &amp;lt;code&amp;gt;IMB&amp;lt;/code&amp;gt; is absent from &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; and has no watch.&lt;br /&gt;
Although an &amp;lt;code&amp;gt;IMA&amp;lt;/code&amp;gt; B watch exists on the same date, that does not prove the two&lt;br /&gt;
identifiers name the same individual.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT DISTINCT quote_literal(os_fol_b_focal_animid) AS source_focal,&lt;br /&gt;
       array_agg(DISTINCT biography_data.animid&lt;br /&gt;
                 ORDER BY biography_data.animid)&lt;br /&gt;
         FILTER (WHERE biography_data.animid IS NOT NULL) AS canonical_matches&lt;br /&gt;
  FROM clean.other_species&lt;br /&gt;
    LEFT JOIN sokwedb.biography_data&lt;br /&gt;
      ON LOWER(biography_data.animid) =&lt;br /&gt;
         LOWER(BTRIM(other_species.os_fol_b_focal_animid))&lt;br /&gt;
  WHERE os_fol_b_focal_animid &amp;lt;&amp;gt; BTRIM(os_fol_b_focal_animid)&lt;br /&gt;
        OR NOT EXISTS (&lt;br /&gt;
             SELECT 1&lt;br /&gt;
               FROM sokwedb.biography_data AS exact_biography&lt;br /&gt;
               WHERE exact_biography.animid =&lt;br /&gt;
                     other_species.os_fol_b_focal_animid)&lt;br /&gt;
  GROUP BY os_fol_b_focal_animid&lt;br /&gt;
  ORDER BY source_focal;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: resolve trimmed values to the&lt;br /&gt;
single case-insensitive canonical &amp;lt;code&amp;gt;BIOGRAPHY_DATA.AnimID&amp;lt;/code&amp;gt;, preserving its&lt;br /&gt;
stored spelling.  This removes the 12 trailing spaces and maps &amp;lt;code&amp;gt;pl&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;PL&amp;lt;/code&amp;gt;.&lt;br /&gt;
Exclude exactly the &amp;lt;code&amp;gt;IMB&amp;lt;/code&amp;gt; row dated 2016-08-27 at 10:30; do not infer &amp;lt;code&amp;gt;IMA&amp;lt;/code&amp;gt;.&lt;br /&gt;
Normalized source keys remain 6,366 unique.  Exclude exactly the &amp;lt;code&amp;gt;IMB&amp;lt;/code&amp;gt; row;&lt;br /&gt;
the independently missing B watch for &amp;lt;code&amp;gt;FU&amp;lt;/code&amp;gt; is handled by Problem #219.&lt;br /&gt;
&lt;br /&gt;
== * (#216) OTHER_SPECIES event times cannot satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Six source rows cannot be represented without inventing or rewriting event&lt;br /&gt;
times.  Three have NULL end times, two have Start after Stop, and one&lt;br /&gt;
additional row starts before 04:00.  One reversed row also starts after 20:00.&lt;br /&gt;
There are no second-bearing times and no Stop values after 20:00.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT os_fol_date, os_fol_b_focal_animid,&lt;br /&gt;
       os_time_begin, os_time_end, os_duration,&lt;br /&gt;
       os_osl_local_species_name&lt;br /&gt;
  FROM clean.other_species&lt;br /&gt;
  WHERE os_time_end IS NULL&lt;br /&gt;
        OR os_time_begin &amp;gt; os_time_end&lt;br /&gt;
        OR os_time_begin &amp;lt; TIME &amp;#039;04:00&amp;#039;&lt;br /&gt;
        OR os_time_begin &amp;gt; TIME &amp;#039;20:00&amp;#039;&lt;br /&gt;
        OR os_time_end &amp;gt; TIME &amp;#039;20:00&amp;#039;&lt;br /&gt;
  ORDER BY os_fol_date, os_fol_b_focal_animid, os_time_begin;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: exclude the exact six-row union&lt;br /&gt;
at the loader boundary.  Do not derive a missing endpoint from Duration or use&lt;br /&gt;
Duration to replace a recorded time.  Preserve both endpoints for every other&lt;br /&gt;
row; &amp;lt;code&amp;gt;OS&amp;lt;/code&amp;gt; events are intervals and cannot use &amp;lt;code&amp;gt;sdb_no_time&amp;lt;/code&amp;gt;.  These six rows&lt;br /&gt;
do not overlap Problems #214, #215, or #219.  Together with Problems #214 and&lt;br /&gt;
#215 they bring the exclusion union to 12 rows.&lt;br /&gt;
&lt;br /&gt;
== (#217) OTHER_SPECIES comments map to event notes ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;EVENTS.Notes&amp;lt;/code&amp;gt; is NOT NULL.  Of 6,366 source rows, 6,352 have NULL&lt;br /&gt;
&amp;lt;code&amp;gt;OS_comments&amp;lt;/code&amp;gt; and 14 have non-NULL comments.  The separate&lt;br /&gt;
&amp;lt;code&amp;gt;OS_local_species_name_written&amp;lt;/code&amp;gt; field contains free-form audit evidence but&lt;br /&gt;
has no production destination and sometimes differs from the classified&lt;br /&gt;
local-species field.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (WHERE os_comments IS NULL) AS null_comments,&lt;br /&gt;
       COUNT(*) FILTER (WHERE os_comments IS NOT NULL) AS recorded_comments,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE os_comments IS NOT NULL&lt;br /&gt;
               AND os_comments &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
               AND BTRIM(os_comments) = &amp;#039;&amp;#039;) AS whitespace_only_comments&lt;br /&gt;
  FROM clean.other_species;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The query returns 6,352 NULL comments, 14 recorded comments, and zero&lt;br /&gt;
nonempty whitespace-only comments.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: map &amp;lt;code&amp;gt;OS_comments&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;EVENTS.Notes&amp;lt;/code&amp;gt;, preserving non-NULL text exactly and mapping SQL NULL to &amp;lt;code&amp;gt;&amp;#039;&amp;#039;&amp;lt;/code&amp;gt;.&lt;br /&gt;
Retain &amp;lt;code&amp;gt;OS_local_species_name_written&amp;lt;/code&amp;gt; in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; as audit evidence; do not&lt;br /&gt;
append it to event notes or use it to override the classified species field.&lt;br /&gt;
This excludes no rows.  Validation must prove two-way parity under&lt;br /&gt;
&amp;lt;code&amp;gt;COALESCE(os_comments, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#218) SPECIES_PRESENT trigger error paths reference nonexistent fields ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Both error-detail paths in&lt;br /&gt;
&amp;lt;code&amp;gt;db/schemas/lib/triggers/create/species_present.m4&amp;lt;/code&amp;gt; reference&lt;br /&gt;
&amp;lt;code&amp;gt;NEW.researchers&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;NEW.nonresearchers&amp;lt;/code&amp;gt;.  Those are &amp;lt;code&amp;gt;HUMANS&amp;lt;/code&amp;gt; columns;&lt;br /&gt;
&amp;lt;code&amp;gt;SPECIES_PRESENT&amp;lt;/code&amp;gt; has only &amp;lt;code&amp;gt;EID&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;Species&amp;lt;/code&amp;gt;.  An invalid related event or a&lt;br /&gt;
conflicting &amp;lt;code&amp;gt;HUMANS&amp;lt;/code&amp;gt; row can therefore obscure the intended integrity error&lt;br /&gt;
with a record-field error.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
grep -nE &amp;#039;NEW\.(researchers|nonresearchers)&amp;#039; \&lt;br /&gt;
  db/schemas/lib/triggers/create/species_present.m4&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The command reports four references, two in each error-detail path.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Resolved on 2026-09-18.  Both invalid attempted-row details now report&lt;br /&gt;
&amp;lt;code&amp;gt;NEW.species&amp;lt;/code&amp;gt;; the related event/watch and conflicting HUMANS details remain.&lt;br /&gt;
The generated trigger was installed in a disposable PostgreSQL 18.6 database,&lt;br /&gt;
and &amp;lt;code&amp;gt;db/tests/species_present_trigger.sql&amp;lt;/code&amp;gt; passed rollback-only tests for wrong&lt;br /&gt;
event behavior, HUMANS conflict, a valid &amp;lt;code&amp;gt;OS&amp;lt;/code&amp;gt; insert, immediate deferred&lt;br /&gt;
constraints, and fixture cleanup.  This is a trigger repair and excludes no&lt;br /&gt;
source rows.&lt;br /&gt;
&lt;br /&gt;
== * (#219) OTHER_SPECIES row has no B watch at the required load stage ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After the normal conversion sequence through &amp;lt;code&amp;gt;load_follow_to_watches&amp;lt;/code&amp;gt;, the&lt;br /&gt;
otherwise valid &amp;lt;code&amp;gt;FU&amp;lt;/code&amp;gt; row dated 2015-12-13 at 16:45 has no B watch.  &amp;lt;code&amp;gt;FU&amp;lt;/code&amp;gt; has no&lt;br /&gt;
source FOLLOW row that day.  Later aggression and grooming loaders create a B&lt;br /&gt;
watch for the same focal/date, but moving OTHER_SPECIES later would make its&lt;br /&gt;
prerequisite depend on unrelated observations and conceal the stage-specific&lt;br /&gt;
source gap.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT os_fol_date, os_fol_b_focal_animid, os_time_begin, os_time_end,&lt;br /&gt;
       os_osl_local_species_name, os_comments&lt;br /&gt;
  FROM clean.other_species&lt;br /&gt;
  WHERE os_fol_date = DATE &amp;#039;2015-12-13&amp;#039;&lt;br /&gt;
        AND os_fol_b_focal_animid = &amp;#039;FU&amp;#039;&lt;br /&gt;
        AND os_time_begin = TIME &amp;#039;16:45&amp;#039;;&lt;br /&gt;
&lt;br /&gt;
SELECT watches.wid&lt;br /&gt;
  FROM sokwedb.watches&lt;br /&gt;
  WHERE watches.animid = &amp;#039;FU&amp;#039;&lt;br /&gt;
        AND watches.date = DATE &amp;#039;2015-12-13&amp;#039;&lt;br /&gt;
        AND watches.type = &amp;#039;B&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At the required post-follow loader stage, the first query returns the one&lt;br /&gt;
&amp;lt;code&amp;gt;VYONDI&amp;lt;/code&amp;gt; interval from 16:45 through 17:00 and the second query returns no row.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: exclude exactly the &amp;lt;code&amp;gt;FU&amp;lt;/code&amp;gt; row dated&lt;br /&gt;
2015-12-13 at 16:45.  Keep OTHER_SPECIES immediately after&lt;br /&gt;
&amp;lt;code&amp;gt;load_follow_to_watches&amp;lt;/code&amp;gt;; do not create a watch and do not reuse a watch that&lt;br /&gt;
only appears after later aggression or grooming loads.  This row overlaps&lt;br /&gt;
none of Problems #214, #215, or #216, increasing the final exclusion union to&lt;br /&gt;
13 rows and leaving 6,353 rows eligible for conversion.&lt;br /&gt;
&lt;br /&gt;
== (#220) FOLLOW_MAP_LOCATION rows can contain two location representations ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Of 1,823,766 &amp;lt;code&amp;gt;FOLLOW_MAP_LOCATION&amp;lt;/code&amp;gt; rows, 703,807 contain both a paper map&lt;br /&gt;
sequence and paired UTM coordinates.  Production stores PAPER and UTM details&lt;br /&gt;
on separate EVENTS rows, so treating the source row as a choice of one detail&lt;br /&gt;
table would discard one recorded representation.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
The read-only &amp;lt;code&amp;gt;conversion/follow_map_location_profile.sql&amp;lt;/code&amp;gt; reports 703,807&lt;br /&gt;
dual rows, 69 paper-only rows, 1,119,890 UTM-only rows, and no row with neither&lt;br /&gt;
representation.  Before exclusions this is 703,876 PAPER representations and&lt;br /&gt;
1,823,697 UTM representations.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: emit one PAPER event and one UTM&lt;br /&gt;
event for every eligible dual row.  Both events use the same location watch&lt;br /&gt;
and source time.  No loader may choose one representation or infer that UTM&lt;br /&gt;
coordinates accompanying a paper map are direct GPS observations.&lt;br /&gt;
&lt;br /&gt;
== * (#221) FOLLOW_MAP_LOCATION focal identities are missing or unsupported ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Location watches require a real &amp;lt;code&amp;gt;BIOGRAPHY_DATA.AnimID&amp;lt;/code&amp;gt;; &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; is prohibited.&lt;br /&gt;
The refreshed source has 6,509 NULL-focal rows and 7,495 rows in 33 trimmed&lt;br /&gt;
labels having no normalized biography match.  The source does not contain&lt;br /&gt;
evidence from which to invent their identities.&lt;br /&gt;
&lt;br /&gt;
Fourteen noncanonical spellings covering 609 rows do resolve uniquely under&lt;br /&gt;
&amp;lt;code&amp;gt;lower(normalize(btrim(value)))&amp;lt;/code&amp;gt;.  They include edge-space variants and the&lt;br /&gt;
case variants &amp;lt;code&amp;gt;Nas&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;NAS&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fd&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FD&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;pl&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;PL&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;sw &amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;SW&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Run &amp;lt;code&amp;gt;conversion/follow_map_location_profile.sql&amp;lt;/code&amp;gt;.  It prints every unmatched&lt;br /&gt;
trimmed label with its row count and date range, and reports 214 raw non-NULL&lt;br /&gt;
spellings, 204 trimmed spellings, and 200 normalized spelling classes.  The&lt;br /&gt;
unmatched labels include punctuation-prefixed labels, apparent suffixed IDs,&lt;br /&gt;
and labels that cannot safely be corrected from spelling alone.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: standardize the 609 uniquely&lt;br /&gt;
matched rows to the stored biography spelling and exclude all 6,509 NULL-focal&lt;br /&gt;
rows and 7,495 unmatched-label rows.  This excludes 14,004 rows before overlap&lt;br /&gt;
with other problems.  Do not derive a focal from coordinates, community, date,&lt;br /&gt;
another observation, or a nearby follow.&lt;br /&gt;
&lt;br /&gt;
== * (#222) FOLLOW_MAP_LOCATION event times cannot be represented ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PAPER and UTM events require a recorded minute from 04:00 through 20:00.  The&lt;br /&gt;
refreshed source has 25,921 NULL times, 676 times before 04:00, and 258 after&lt;br /&gt;
20:00.  One recorded time is midnight, so midnight cannot represent NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/follow_map_location_profile.sql&amp;lt;/code&amp;gt; reports the exact classes and&lt;br /&gt;
representation impact.  NULL times affect 25,921 PAPER and 25,914 UTM&lt;br /&gt;
representations; early times affect 676 UTM representations; late times&lt;br /&gt;
affect 7 PAPER and 258 UTM representations.  Recorded seconds are always zero.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: exclude all 25,921 NULL-time rows,&lt;br /&gt;
676 rows before 04:00, and 258 rows after 20:00.  This excludes 26,855 rows&lt;br /&gt;
before overlap with other problems.  Do not map NULL to midnight, clamp&lt;br /&gt;
out-of-range times, round times, or derive a time from neighboring source rows.&lt;br /&gt;
&lt;br /&gt;
== * (#223) FOLLOW_MAP_LOCATION event keys collide ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production allows at most one PAPER and one UTM event for a location watch at&lt;br /&gt;
a given time.  Before focal normalization and after omitting NULL times, the&lt;br /&gt;
source has 23,801 colliding PAPER focal/date/time keys and 25,561 colliding UTM&lt;br /&gt;
keys.  Generated IDs cannot make those keys representable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
The profile reports 47,603 PAPER rows in colliding keys, with 23,802 excess&lt;br /&gt;
rows, and 51,255 UTM rows, with 25,694 excess rows.  Only 75 PAPER keys (150&lt;br /&gt;
rows) and 150 UTM keys (300 rows) have one identical representation value;&lt;br /&gt;
23,726 PAPER keys and 25,411 UTM keys contain genuinely different recorded&lt;br /&gt;
representations.  One UTM key has 14 source rows.&lt;br /&gt;
&lt;br /&gt;
The complete source additionally has 151 exact duplicate pairs across all 14&lt;br /&gt;
columns.  A generated conversion ID does not establish that either copy is a&lt;br /&gt;
separate production observation.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: exclude every source row whose&lt;br /&gt;
canonical focal/date/time key collides for either available representation.&lt;br /&gt;
This excludes the entire source row, rather than retaining half of a dual row,&lt;br /&gt;
and removes 51,457 rows before overlap with other problems.  Do not add&lt;br /&gt;
arbitrary time offsets, select a row by generated order, or let &amp;lt;code&amp;gt;ON CONFLICT&amp;lt;/code&amp;gt;&lt;br /&gt;
discard rows.&lt;br /&gt;
&lt;br /&gt;
== * (#224) FOLLOW_MAP_LOCATION community recovery counts depend on focal normalization ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The investigator approved dated &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; as canonical, &amp;lt;code&amp;gt;Unknown&amp;lt;/code&amp;gt; only for&lt;br /&gt;
a NULL source community without dated membership, and exclusion of unsupported&lt;br /&gt;
non-NULL communities and valid non-NULL communities without matching dated&lt;br /&gt;
membership or disagreeing with membership.&lt;br /&gt;
&lt;br /&gt;
The previously approved exact counts used a case-sensitive&lt;br /&gt;
&amp;lt;code&amp;gt;btrim(focal) = COMM_MEMBS.AnimID&amp;lt;/code&amp;gt; join.  The loader specification instead&lt;br /&gt;
requires canonical focal matching by &amp;lt;code&amp;gt;lower(normalize(btrim(value)))&amp;lt;/code&amp;gt;.&lt;br /&gt;
Applying that rule resolves 118 additional rows through &amp;lt;code&amp;gt;Nas&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;NAS&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fd&amp;lt;/code&amp;gt;&lt;br /&gt;
to &amp;lt;code&amp;gt;FD&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;pl&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;PL&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;sw &amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;SW&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
On the refreshed source, canonical matching reports:&lt;br /&gt;
&lt;br /&gt;
* 1,095,525 NULL-source rows with membership;&lt;br /&gt;
* 14,412 NULL-source rows without membership;&lt;br /&gt;
* 712,457 valid-source rows agreeing with membership;&lt;br /&gt;
* 245 valid-source rows without membership;&lt;br /&gt;
* 1,105 valid-source rows disagreeing with membership; and&lt;br /&gt;
* 22 unsupported non-NULL source communities.&lt;br /&gt;
&lt;br /&gt;
The canonical result would exclude 1,372 rows rather than 1,455.  The old&lt;br /&gt;
trim-only result remains exactly 1,095,490, 14,447, 712,374, 328, 1,105, and&lt;br /&gt;
22 respectively.  The 83-row exclusion difference is 51 &amp;lt;code&amp;gt;fd&amp;lt;/code&amp;gt;, 14 &amp;lt;code&amp;gt;pl&amp;lt;/code&amp;gt;, and 18&lt;br /&gt;
&amp;lt;code&amp;gt;sw &amp;lt;/code&amp;gt; rows whose valid source community agrees with canonical membership.  The&lt;br /&gt;
other 35 newly resolved rows are NULL-community &amp;lt;code&amp;gt;Nas&amp;lt;/code&amp;gt; rows and would change&lt;br /&gt;
from &amp;lt;code&amp;gt;Unknown&amp;lt;/code&amp;gt; to membership community &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Canonical focal standardization was approved under Problem #221.  The final&lt;br /&gt;
Problem #224 exclusion union is therefore 1,372 rows: 22 unsupported non-NULL&lt;br /&gt;
communities, 245 valid communities without dated membership, and 1,105 valid&lt;br /&gt;
communities disagreeing with membership.  Do not preserve the old counts by&lt;br /&gt;
using a less canonical identity join.&lt;br /&gt;
&lt;br /&gt;
== (#225) LOCATION_ORIGINS lacks approved support values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;LOCATION_ORIGINS&amp;lt;/code&amp;gt; is empty, but every PAPER and UTM detail requires an origin.&lt;br /&gt;
The source contains &amp;lt;code&amp;gt;GPS&amp;lt;/code&amp;gt; on 1,113,996 rows, &amp;lt;code&amp;gt;Paper Map&amp;lt;/code&amp;gt; on 15,165 rows,&lt;br /&gt;
&amp;lt;code&amp;gt;paper maps&amp;lt;/code&amp;gt; on 27,273 rows, and NULL on 667,332 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
The profile reports that NULL origin affects 661,438 paper representations and&lt;br /&gt;
667,332 UTM representations.  Of the latter, 5,894 are UTM-only rows.  A UTM&lt;br /&gt;
representation therefore does not prove GPS origin, and dual historical rows&lt;br /&gt;
may contain coordinates converted from a paper map.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18.  Populate canonical origins &amp;lt;code&amp;gt;GPS&amp;lt;/code&amp;gt;&lt;br /&gt;
(&amp;lt;code&amp;gt;GPS&amp;lt;/code&amp;gt;), &amp;lt;code&amp;gt;PAPER&amp;lt;/code&amp;gt; (&amp;lt;code&amp;gt;Paper map&amp;lt;/code&amp;gt;), and &amp;lt;code&amp;gt;UNKNOWN&amp;lt;/code&amp;gt; (&amp;lt;code&amp;gt;Unknown&amp;lt;/code&amp;gt;).  The uppercase code&lt;br /&gt;
is required by the production support-table constraint.  Map source &amp;lt;code&amp;gt;GPS&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;GPS&amp;lt;/code&amp;gt;, both &amp;lt;code&amp;gt;Paper Map&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;paper maps&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;PAPER&amp;lt;/code&amp;gt;, and SQL NULL to &amp;lt;code&amp;gt;UNKNOWN&amp;lt;/code&amp;gt;.&lt;br /&gt;
Apply the same source mapping to both representations; do not infer GPS solely&lt;br /&gt;
from the presence of coordinates.&lt;br /&gt;
&lt;br /&gt;
== (#226) FOLLOW_MAP_LOCATION entered-person support is incomplete ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;Entered&amp;lt;/code&amp;gt; is NULL on 1,781,328 rows.  Recorded values include two labels for&lt;br /&gt;
Danielle Lodge and the full name &amp;lt;code&amp;gt;Julianna Turner&amp;lt;/code&amp;gt;; those values did not&lt;br /&gt;
previously have complete PEOPLE support.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
The refreshed profile reports 17,794 &amp;lt;code&amp;gt;DL&amp;lt;/code&amp;gt;, 8,902 &amp;lt;code&amp;gt;Danielle Lodge&amp;lt;/code&amp;gt;, 7,968&lt;br /&gt;
&amp;lt;code&amp;gt;Julianna Turner&amp;lt;/code&amp;gt;, 2,314 &amp;lt;code&amp;gt;KS&amp;lt;/code&amp;gt;, 4,806 &amp;lt;code&amp;gt;SM&amp;lt;/code&amp;gt;, and 654 &amp;lt;code&amp;gt;SR&amp;lt;/code&amp;gt; rows.  The refreshed&lt;br /&gt;
support table contains &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;KS&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;SM&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;SR&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;DL&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;Julianna Turner&amp;lt;/code&amp;gt;&lt;br /&gt;
with no normalized duplicate.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18.  Map NULL to &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;; map both &amp;lt;code&amp;gt;DL&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;Danielle Lodge&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;DL&amp;lt;/code&amp;gt;; map &amp;lt;code&amp;gt;Julianna Turner&amp;lt;/code&amp;gt; to that same-named code;&lt;br /&gt;
and preserve &amp;lt;code&amp;gt;KS&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;SM&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;SR&amp;lt;/code&amp;gt;.  Match with&lt;br /&gt;
&amp;lt;code&amp;gt;lower(normalize(btrim(value)))&amp;lt;/code&amp;gt;.  The manually maintained &amp;lt;code&amp;gt;clean.people&amp;lt;/code&amp;gt; seed&lt;br /&gt;
contains &amp;lt;code&amp;gt;DL&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;Julianna Turner&amp;lt;/code&amp;gt;, both described as follow map location data&lt;br /&gt;
entry.  This decision excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#227) FOLLOW_MAP_LOCATION contains spatial anomalies ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Thirty-nine UTM rows have coordinates outside the expected local orientation,&lt;br /&gt;
and two rows have elevation 20,000.  The production numeric limits equal the&lt;br /&gt;
source extrema, so constraint acceptance is not evidence of geographic&lt;br /&gt;
validity.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/follow_map_location_profile.sql&amp;lt;/code&amp;gt; prints all 41 anomaly rows.  The&lt;br /&gt;
39 coordinate rows include equal axes, axes both near 9.48 million, and values&lt;br /&gt;
that appear to have missing digits.  The two elevation rows are dated&lt;br /&gt;
2020-03-08 (&amp;lt;code&amp;gt;KOM&amp;lt;/code&amp;gt;) and 2021-06-06 (&amp;lt;code&amp;gt;AP&amp;lt;/code&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: preserve all 39 unusual coordinate&lt;br /&gt;
rows and both elevation-20,000 rows exactly as recorded.  They satisfy current&lt;br /&gt;
production bounds.  Do not swap axes, add digits, replace elevation by&lt;br /&gt;
heuristic, or exclude these rows solely because they are anomalous.&lt;br /&gt;
&lt;br /&gt;
== (#228) FOLLOW_MAP_LOCATION has NULL optional metadata ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;FollowNum&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;Notes&amp;lt;/code&amp;gt; are NOT NULL but allow empty text.  Source&lt;br /&gt;
&amp;lt;code&amp;gt;FollowNum&amp;lt;/code&amp;gt; is NULL on 1,163,270 rows and &amp;lt;code&amp;gt;Notes&amp;lt;/code&amp;gt; is NULL on 1,823,410 rows.&lt;br /&gt;
&amp;lt;code&amp;gt;fml_update&amp;lt;/code&amp;gt; is populated on 48,796 rows but has no production destination.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
The profile reports 18,541 distinct populated follow numbers of length 5&lt;br /&gt;
through 10, and 356 populated notes containing 197 distinct values.  Every&lt;br /&gt;
populated value can be preserved exactly.  Every populated &amp;lt;code&amp;gt;fml_update&amp;lt;/code&amp;gt; is&lt;br /&gt;
2018-05-20.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: map SQL NULL &amp;lt;code&amp;gt;FollowNum&amp;lt;/code&amp;gt; and&lt;br /&gt;
&amp;lt;code&amp;gt;Notes&amp;lt;/code&amp;gt; to empty text while preserving populated values exactly.  Retain&lt;br /&gt;
&amp;lt;code&amp;gt;fml_update&amp;lt;/code&amp;gt; only in the clean audit source and do not append it to notes.&lt;br /&gt;
&lt;br /&gt;
== * (#229) FOLLOW_MAP_LOCATION dates exceed focal departure dates ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;WATCHES&amp;lt;/code&amp;gt; rejects a watch before the focal&amp;#039;s birth or entry and after the&lt;br /&gt;
focal&amp;#039;s departure from study.  The source has 229 canonically matched rows&lt;br /&gt;
after &amp;lt;code&amp;gt;BIOGRAPHY_DATA.DepartDate&amp;lt;/code&amp;gt;, of which 47 are already excluded by both&lt;br /&gt;
Problems #222 and #224.  The approved Problems #220-#228 projection therefore&lt;br /&gt;
still contained 182 rows after departure.  There are no projected rows before&lt;br /&gt;
birth or entry.&lt;br /&gt;
&lt;br /&gt;
All 182 affected rows are UTM-only GPS records with NULL source community,&lt;br /&gt;
NULL entered person, and valid event times.  They form two location watches:&lt;br /&gt;
&lt;br /&gt;
* focal &amp;lt;code&amp;gt;S&amp;lt;/code&amp;gt;, 134 rows on 2017-11-29 from 06:46 through 18:42, after departure&lt;br /&gt;
  date 1968-01-19; and&lt;br /&gt;
* focal &amp;lt;code&amp;gt;DE&amp;lt;/code&amp;gt;, 48 rows on 2017-12-27 from 12:41 through 17:02, after departure&lt;br /&gt;
  date 1974-05-07.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT source.*,&lt;br /&gt;
       biography_data.birthdate,&lt;br /&gt;
       biography_data.entrydate,&lt;br /&gt;
       biography_data.departdate&lt;br /&gt;
  FROM clean.follow_map_location AS source&lt;br /&gt;
  JOIN sokwedb.biography_data&lt;br /&gt;
    ON lower(normalize(biography_data.animid)) =&lt;br /&gt;
         lower(normalize(btrim(source.fml_fol_b_focal_animid)))&lt;br /&gt;
 WHERE source.fml_fol_date &amp;lt; biography_data.birthdate&lt;br /&gt;
       OR source.fml_fol_date &amp;lt; biography_data.entrydate&lt;br /&gt;
       OR source.fml_fol_date &amp;gt; biography_data.departdate&lt;br /&gt;
 ORDER BY source.fml_conversion_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The first full chunked load stopped when inserting the &amp;lt;code&amp;gt;S&amp;lt;/code&amp;gt; location watch on&lt;br /&gt;
2017-11-29 because the &amp;lt;code&amp;gt;WATCHES&amp;lt;/code&amp;gt; trigger reported its 1968-01-19 departure.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: exclude all post-departure rows.&lt;br /&gt;
The predicate identifies 229 raw rows; 47 overlap both Problems #222 and #224,&lt;br /&gt;
so Problem #229 adds exactly 182 rows to the exclusion union.  The final&lt;br /&gt;
distinct union is 93,470, leaving 1,730,296 eligible source rows.  Do not weaken&lt;br /&gt;
the production biography-date trigger, substitute another focal, or alter the&lt;br /&gt;
recorded location dates.&lt;br /&gt;
&lt;br /&gt;
== (#230) Location events lack a matching focal arrival interval ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production documentation expects location times to fall within a focal&amp;#039;s&lt;br /&gt;
arrival interval on a B watch, but the database reports this condition as a&lt;br /&gt;
warning rather than rejecting the event.  Of 2,358,936 loaded PAPER and UTM&lt;br /&gt;
events, 779,682 do not fall within a matching focal arrival interval.&lt;br /&gt;
&lt;br /&gt;
The refreshed post-load profile separates them as follows:&lt;br /&gt;
&lt;br /&gt;
* 724,075 events have no B watch for the focal/date: 196 PAPER and 723,879 UTM;&lt;br /&gt;
* 404 events have a B watch but no ARR interval: 202 PAPER and 202 UTM; and&lt;br /&gt;
* 55,203 events are outside every ARR interval on an existing B watch: 19,118&lt;br /&gt;
  PAPER and 36,085 UTM.&lt;br /&gt;
&lt;br /&gt;
These classes cover 1973-01-04 through 2022-11-28.  The large no-B-watch UTM&lt;br /&gt;
class reflects location coverage continuing well beyond the consolidated&lt;br /&gt;
follow data; it is not evidence from which to invent a follow or arrival time.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Join location &amp;lt;code&amp;gt;EVENTS&amp;lt;/code&amp;gt; and their &amp;lt;code&amp;gt;L&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;WATCHES&amp;lt;/code&amp;gt; rows to B watches on focal/date,&lt;br /&gt;
then test whether the location start falls between an ARR event&amp;#039;s start and&lt;br /&gt;
stop.  &amp;lt;code&amp;gt;conversion/load_follow_map_location_validation.sql&amp;lt;/code&amp;gt; reports the exact&lt;br /&gt;
779,682-event total after proving source-to-target parity.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: preserve all 779,682 events as&lt;br /&gt;
recorded.  The production schema permits them and the arrival relationship is&lt;br /&gt;
documentation guidance enforced as a warning.  Validation asserts the exact&lt;br /&gt;
count and continues to report it.  Do not create B watches, manufacture ARR&lt;br /&gt;
intervals, alter location times, or exclude otherwise eligible source rows to&lt;br /&gt;
satisfy this guidance.&lt;br /&gt;
&lt;br /&gt;
== * (#231) CHIMPS_ON_TIKIS source identifiers lack biography support ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Three &amp;lt;code&amp;gt;clean.chimps_on_tikis_lookup&amp;lt;/code&amp;gt; rows cannot satisfy the required&lt;br /&gt;
&amp;lt;code&amp;gt;CHIMPS_ON_TIKIS.AnimID&amp;lt;/code&amp;gt; foreign key: &amp;lt;code&amp;gt;MGENI&amp;lt;/code&amp;gt; occurs once in &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt; and once in&lt;br /&gt;
&amp;lt;code&amp;gt;MT&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;VAN mtoto&amp;lt;/code&amp;gt; occurs once in &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt;.  &amp;lt;code&amp;gt;MGENI&amp;lt;/code&amp;gt; does not identify a&lt;br /&gt;
particular production individual.  &amp;lt;code&amp;gt;VAN mtoto&amp;lt;/code&amp;gt; is not a production AnimID.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;VIS&amp;lt;/code&amp;gt; is VAN&amp;#039;s only supported child alive on the Tiki creation and deployment&lt;br /&gt;
dates and had &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt; membership, but that evidence does not establish that the&lt;br /&gt;
source label denotes &amp;lt;code&amp;gt;VIS&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT source.*&lt;br /&gt;
  FROM clean.chimps_on_tikis_lookup AS source&lt;br /&gt;
  LEFT JOIN sokwedb.biography_data AS biography&lt;br /&gt;
    ON biography.animid = source.b_animid&lt;br /&gt;
  WHERE biography.animid IS NULL&lt;br /&gt;
  ORDER BY source.tiki_creation_date,&lt;br /&gt;
           source.community_id,&lt;br /&gt;
           source.b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns exactly:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
 b_animid  | tiki_creation_date | community_id | tiki_deployment_date | new_addition&lt;br /&gt;
-----------+--------------------+--------------+----------------------+-------------&lt;br /&gt;
 MGENI     | 2020-08-04         | KK           | 2020-08-20           | N&lt;br /&gt;
 VAN mtoto | 2020-08-04         | KK           | 2020-08-20           | Y&lt;br /&gt;
 MGENI     | 2020-08-04         | MT           | 2020-08-20           | N&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: exclude exactly these three&lt;br /&gt;
complete source tuples.  Do not create a biography row for &amp;lt;code&amp;gt;MGENI&amp;lt;/code&amp;gt;, infer an&lt;br /&gt;
individual for either &amp;lt;code&amp;gt;MGENI&amp;lt;/code&amp;gt; row, or map &amp;lt;code&amp;gt;VAN mtoto&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;VIS&amp;lt;/code&amp;gt;.  The loader&lt;br /&gt;
sanity check must prove the exact exclusion inventory and reject any remaining&lt;br /&gt;
unsupported identifier.  The dated source has 84 rows, leaving 81 eligible&lt;br /&gt;
rows.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  The shared projection compares all&lt;br /&gt;
five source columns with &amp;lt;code&amp;gt;IS NOT DISTINCT FROM&amp;lt;/code&amp;gt; and excludes only the approved&lt;br /&gt;
tuples.  Read-only sanity reproduced 84 source rows, the exact three-row&lt;br /&gt;
exclusion multiset, and 81 supported unique projected rows.  The direct load&lt;br /&gt;
produced 81 target rows, including 13 &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt; and 68 &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt; newcomer values;&lt;br /&gt;
bidirectional &amp;lt;code&amp;gt;EXCEPT ALL&amp;lt;/code&amp;gt; returned no payload differences.  An intentional&lt;br /&gt;
post-insert failure rolled all 81 rows back to zero.  An approved clean retry&lt;br /&gt;
reproduced the same payload digest, &amp;lt;code&amp;gt;28e17175eba7622cbf9afc7130dc1900&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#232) CHIMPS_ON_TIKIS deployment-date index targets AnimID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The index named &amp;lt;code&amp;gt;chimps_on_tikis_tikideployementdate&amp;lt;/code&amp;gt; is defined on &amp;lt;code&amp;gt;animid&amp;lt;/code&amp;gt;&lt;br /&gt;
instead of &amp;lt;code&amp;gt;tikideploymentdate&amp;lt;/code&amp;gt;.  It duplicates the separate AnimID index and&lt;br /&gt;
leaves the documented deployment-date lookup unindexed.  The index name also&lt;br /&gt;
misspells &amp;quot;deployment&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT indexname, indexdef&lt;br /&gt;
  FROM pg_indexes&lt;br /&gt;
  WHERE schemaname = &amp;#039;housekeeping&amp;#039;&lt;br /&gt;
        AND tablename = &amp;#039;chimps_on_tikis&amp;#039;&lt;br /&gt;
        AND indexname = &amp;#039;chimps_on_tikis_tikideployementdate&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The live PostgreSQL 18.6 database reports an index using &amp;lt;code&amp;gt;btree (animid)&amp;lt;/code&amp;gt;.&lt;br /&gt;
The owning definition is in&lt;br /&gt;
&amp;lt;code&amp;gt;db/schemas/housekeeping/indexes/create/chimps_on_tikis.m4&amp;lt;/code&amp;gt;; the corresponding&lt;br /&gt;
drop name is in the neighboring drop M4.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: before loading this table, replace&lt;br /&gt;
the defective index with correctly spelled&lt;br /&gt;
&amp;lt;code&amp;gt;chimps_on_tikis_tikideploymentdate&amp;lt;/code&amp;gt; on &amp;lt;code&amp;gt;(tikideploymentdate)&amp;lt;/code&amp;gt;.  Update both&lt;br /&gt;
owning M4 files, regenerate owned index SQL artifacts, and verify the live&lt;br /&gt;
index definition in a disposable database.  This schema prerequisite excludes&lt;br /&gt;
no source rows.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  The owning create and drop M4 files&lt;br /&gt;
now use the corrected index, while the drop definition also removes the legacy&lt;br /&gt;
misspelled name during an in-place migration.  Repository generation rebuilt&lt;br /&gt;
the per-object and aggregate index SQL.  Both the focused local migration and&lt;br /&gt;
the freshly rebuilt disposable integration schema reported exactly one index&lt;br /&gt;
on &amp;lt;code&amp;gt;tikideploymentdate&amp;lt;/code&amp;gt;, named&lt;br /&gt;
&amp;lt;code&amp;gt;chimps_on_tikis_tikideploymentdate&amp;lt;/code&amp;gt;, and no legacy index.&lt;br /&gt;
&lt;br /&gt;
== * (#233) SIV_STATUS_BOUT comma-bearing names shifted into later columns ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Five source rows have a comma-bearing original name split across&lt;br /&gt;
&amp;lt;code&amp;gt;SSB_names_used&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;SSB_in_bio_table&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;SSB_notes&amp;lt;/code&amp;gt;.  The leading name&lt;br /&gt;
fragment is in &amp;lt;code&amp;gt;SSB_names_used&amp;lt;/code&amp;gt;, the trailing fragment is in&lt;br /&gt;
&amp;lt;code&amp;gt;SSB_in_bio_table&amp;lt;/code&amp;gt;, and the actual biography flag &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; is in &amp;lt;code&amp;gt;SSB_notes&amp;lt;/code&amp;gt;.&lt;br /&gt;
These malformed values already occur in &amp;lt;code&amp;gt;easy.siv_status_bout&amp;lt;/code&amp;gt;; they were not&lt;br /&gt;
introduced by a clean-stage transformation.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT source.*&lt;br /&gt;
  FROM clean.siv_status_bout AS source&lt;br /&gt;
  WHERE source.ssb_in_bio_table NOT IN (&amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;)&lt;br /&gt;
  ORDER BY source.ssb_b_animid, source.ssb_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns exactly five rows: &amp;lt;code&amp;gt;DB&amp;lt;/code&amp;gt; with fragments &amp;lt;code&amp;gt;&amp;quot;Darbee&amp;lt;/code&amp;gt; and&lt;br /&gt;
&amp;lt;code&amp;gt;NA&amp;quot;&amp;lt;/code&amp;gt;; &amp;lt;code&amp;gt;TOF&amp;lt;/code&amp;gt; with &amp;lt;code&amp;gt;&amp;quot;Tofiki&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;gone&amp;quot;&amp;lt;/code&amp;gt;; and all three &amp;lt;code&amp;gt;TT&amp;lt;/code&amp;gt; bouts with&lt;br /&gt;
&amp;lt;code&amp;gt;&amp;quot;Kati (Tita)&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;Kati-tita&amp;quot;&amp;lt;/code&amp;gt;.  Each row has &amp;lt;code&amp;gt;SSB_notes = &amp;#039;1&amp;#039;&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: exclude these five exact complete&lt;br /&gt;
source tuples.  The final decision supersedes an earlier consideration of&lt;br /&gt;
reconstructing the values in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt;.  Do not repair the imported fields,&lt;br /&gt;
infer punctuation, or use a broad identifier-only exclusion.  Sanity must&lt;br /&gt;
prove the exact five-row multiset, and the canonical projection must compare&lt;br /&gt;
all ten source columns with NULL-safe equality.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  &amp;lt;code&amp;gt;conversion/siv_status_bout_projection.sql&amp;lt;/code&amp;gt;&lt;br /&gt;
encodes the exact five-row &amp;lt;code&amp;gt;DB&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;TOF&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;TT&amp;lt;/code&amp;gt; tuple set, including the two&lt;br /&gt;
literal embedded double-quote characters recovered with &amp;lt;code&amp;gt;quote_literal&amp;lt;/code&amp;gt;.&lt;br /&gt;
Read-only sanity, run against the local PostgreSQL 18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; database,&lt;br /&gt;
reproduced this exact multiset with bidirectional &amp;lt;code&amp;gt;EXCEPT ALL&amp;lt;/code&amp;gt; and confirmed&lt;br /&gt;
it overlaps Problem #234 in exactly one row and Problem #235 in none.&lt;br /&gt;
&lt;br /&gt;
== * (#234) SIV_STATUS_BOUT required test dates are missing ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Twelve source rows cannot satisfy the target&amp;#039;s non-NULL &amp;lt;code&amp;gt;FirstTestDate&amp;lt;/code&amp;gt; and&lt;br /&gt;
&amp;lt;code&amp;gt;LastTestDate&amp;lt;/code&amp;gt; columns.  All 11 &amp;lt;code&amp;gt;unk&amp;lt;/code&amp;gt; bouts have both dates absent.  The &amp;lt;code&amp;gt;VID&amp;lt;/code&amp;gt;&lt;br /&gt;
positive bout has &amp;lt;code&amp;gt;FirstTestDate = 2000-09-09&amp;lt;/code&amp;gt; and a NULL &amp;lt;code&amp;gt;LastTestDate&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT source.*&lt;br /&gt;
  FROM clean.siv_status_bout AS source&lt;br /&gt;
  WHERE source.ssb_first_test_date IS NULL&lt;br /&gt;
     OR source.ssb_last_test_date IS NULL&lt;br /&gt;
  ORDER BY source.ssb_b_animid, source.ssb_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 11 unknown bouts belong to &amp;lt;code&amp;gt;BIM&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;CH100&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;ERI&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GIM&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;HAI&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;MAK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;RUD&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;TT&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;YD&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;ZS&amp;lt;/code&amp;gt;.  The &amp;lt;code&amp;gt;TT&amp;lt;/code&amp;gt; unknown bout is also one of the&lt;br /&gt;
five Problem #233 shifted rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: exclude the 12 exact complete&lt;br /&gt;
source tuples.  Do not fill absent tests from bout boundaries and do not change&lt;br /&gt;
target nullability.  Match all ten source columns with NULL-safe equality.&lt;br /&gt;
Because the &amp;lt;code&amp;gt;TT&amp;lt;/code&amp;gt; row overlaps Problem #233, count it once in the combined&lt;br /&gt;
exclusion set.&lt;br /&gt;
&lt;br /&gt;
This decision removes every source &amp;lt;code&amp;gt;unk&amp;lt;/code&amp;gt; bout.  Ten retained positive rows&lt;br /&gt;
therefore keep &amp;lt;code&amp;gt;BoutNumber = 3&amp;lt;/code&amp;gt; without a loaded bout 2.  Preserve those source&lt;br /&gt;
bout numbers; do not renumber or synthesize replacement bouts.  Update the&lt;br /&gt;
production documentation to report the resulting snapshot accurately.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  The projection excludes the exact&lt;br /&gt;
12-row tuple set, including the &amp;lt;code&amp;gt;TT&amp;lt;/code&amp;gt; row shared with Problem #233.  Sanity&lt;br /&gt;
reproduced 11 NULL &amp;lt;code&amp;gt;FirstTestDate&amp;lt;/code&amp;gt; rows, 12 NULL &amp;lt;code&amp;gt;LastTestDate&amp;lt;/code&amp;gt; rows, and&lt;br /&gt;
confirmed the resulting 164-row projection contains zero &amp;lt;code&amp;gt;unk&amp;lt;/code&amp;gt; rows and&lt;br /&gt;
exactly 10 &amp;lt;code&amp;gt;BoutNumber = 3&amp;lt;/code&amp;gt; rows with no corresponding loaded&lt;br /&gt;
&amp;lt;code&amp;gt;BoutNumber = 2&amp;lt;/code&amp;gt; row, including &amp;lt;code&amp;gt;CH100&amp;lt;/code&amp;gt;, whose eligible negative and&lt;br /&gt;
positive bouts are retained.&lt;br /&gt;
&lt;br /&gt;
== * (#235) SIV_STATUS_BOUT AMA incorrectly claims biography support ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The source row for &amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt; has &amp;lt;code&amp;gt;SSB_in_bio_table = &amp;#039;1&amp;#039;&amp;lt;/code&amp;gt;, but no exact &amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt;&lt;br /&gt;
identifier exists in &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt;.  &amp;lt;code&amp;gt;AME&amp;lt;/code&amp;gt; exists and the source note says&lt;br /&gt;
the row includes &amp;lt;code&amp;gt;AME&amp;lt;/code&amp;gt; samples, but that is not sufficient evidence to replace&lt;br /&gt;
the source identifier.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT source.*&lt;br /&gt;
  FROM clean.siv_status_bout AS source&lt;br /&gt;
  WHERE source.ssb_b_animid = &amp;#039;AMA&amp;#039;&lt;br /&gt;
    AND source.ssb_in_bio_table = &amp;#039;1&amp;#039;&lt;br /&gt;
    AND NOT EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
            FROM sokwedb.biography_data AS biography&lt;br /&gt;
            WHERE biography.animid = source.ssb_b_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns one &amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt; negative bout from 2014-03-20 through&lt;br /&gt;
2019-04-05, with tests from 2017-01-31 through 2019-04-05 and bout number 1.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: exclude this exact complete source&lt;br /&gt;
tuple.  Do not map &amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;AME&amp;lt;/code&amp;gt;, recompute the flag, or create biography&lt;br /&gt;
support.  Preserve all other unsupported source identifiers whose flags are&lt;br /&gt;
&amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, as the target intentionally permits identifiers outside biography.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  The projection excludes the exact&lt;br /&gt;
&amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt; tuple, including its literal doubled-double-quote note text.  Sanity&lt;br /&gt;
confirmed the 21 remaining unsupported identifiers (&amp;lt;code&amp;gt;CH064&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;CH128&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;TUM&amp;lt;/code&amp;gt;) are retained with &amp;lt;code&amp;gt;InBioTable = FALSE&amp;lt;/code&amp;gt; and that no eligible row&lt;br /&gt;
has an &amp;lt;code&amp;gt;InBioTable&amp;lt;/code&amp;gt;/biography-support mismatch.&lt;br /&gt;
&lt;br /&gt;
== (#236) SIV_STATUS_BOUT absent notes cannot satisfy the target ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The source contains 169 NULL &amp;lt;code&amp;gt;SSB_notes&amp;lt;/code&amp;gt; values, while target &amp;lt;code&amp;gt;Notes&amp;lt;/code&amp;gt; is&lt;br /&gt;
non-NULL.  The target&amp;#039;s &amp;lt;code&amp;gt;notonlyspaces_check&amp;lt;/code&amp;gt; explicitly permits the empty&lt;br /&gt;
string as the representation of no note.&lt;br /&gt;
&lt;br /&gt;
After the approved Problems #233 through #235 exclusion union, 158 eligible&lt;br /&gt;
rows have NULL notes and six have non-NULL notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*)&lt;br /&gt;
  FROM clean.siv_status_bout AS source&lt;br /&gt;
  WHERE source.ssb_notes IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns 169.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: map an eligible NULL note to the&lt;br /&gt;
empty string with &amp;lt;code&amp;gt;COALESCE(source.ssb_notes, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt;.  Preserve every non-NULL&lt;br /&gt;
note exactly.  Sanity and validation must prove that this is the only notes&lt;br /&gt;
transformation and that no projected note contains only spaces.&lt;br /&gt;
&lt;br /&gt;
Problems #233, #234, and #235 contain 5, 12, and 1 rows respectively.  One&lt;br /&gt;
&amp;lt;code&amp;gt;TT&amp;lt;/code&amp;gt; row overlaps #233 and #234, so their union excludes 17 distinct rows from&lt;br /&gt;
181 and leaves 164 rows for conversion.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  &amp;lt;code&amp;gt;COALESCE(source.ssb_notes, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt; is&lt;br /&gt;
the projection&amp;#039;s only notes transformation.  Read-only sanity and validation&lt;br /&gt;
against the local PostgreSQL 18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; database confirmed 158 loaded&lt;br /&gt;
empty notes came only from NULL source notes, all 6 non-NULL notes are&lt;br /&gt;
preserved byte-for-byte, and no loaded note contains only spaces.&lt;br /&gt;
&lt;br /&gt;
The direct set-based loader&lt;br /&gt;
(&amp;lt;code&amp;gt;conversion/load_siv_status_bout.sql&amp;lt;/code&amp;gt;) inserted exactly 164 rows: 132 &amp;lt;code&amp;gt;neg&amp;lt;/code&amp;gt;,&lt;br /&gt;
32 &amp;lt;code&amp;gt;pos&amp;lt;/code&amp;gt;, 0 &amp;lt;code&amp;gt;unk&amp;lt;/code&amp;gt;; 154 &amp;lt;code&amp;gt;BoutNumber = 1&amp;lt;/code&amp;gt;, 0 &amp;lt;code&amp;gt;BoutNumber = 2&amp;lt;/code&amp;gt;, 10&lt;br /&gt;
&amp;lt;code&amp;gt;BoutNumber = 3&amp;lt;/code&amp;gt;; 142 &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt; and 22 &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;InBioTable&amp;lt;/code&amp;gt; values; and 158&lt;br /&gt;
empty/6 nonempty notes.  Bidirectional &amp;lt;code&amp;gt;EXCEPT ALL&amp;lt;/code&amp;gt; between the canonical&lt;br /&gt;
projection and the loaded rows returned no differences in either direction.&lt;br /&gt;
An intentional post-insert failure (&amp;lt;code&amp;gt;SELECT 1/0&amp;lt;/code&amp;gt;) rolled the single&lt;br /&gt;
transaction back to zero target rows.  An authorized &amp;lt;code&amp;gt;TRUNCATE&amp;lt;/code&amp;gt; of only the&lt;br /&gt;
disposable target table followed by an unchanged sanity/load/validation&lt;br /&gt;
sequence reproduced the identical sorted ten-column payload digest,&lt;br /&gt;
&amp;lt;code&amp;gt;8b243c9b7f53d79f9fa1eae640e26b51&amp;lt;/code&amp;gt;, on both runs.  &amp;lt;code&amp;gt;git diff --check&amp;lt;/code&amp;gt;, a&lt;br /&gt;
local-only Make dry run of the full &amp;lt;code&amp;gt;load_data&amp;lt;/code&amp;gt; sequence, and a direct &amp;lt;code&amp;gt;m4&amp;lt;/code&amp;gt;&lt;br /&gt;
and repository doc-build render of &amp;lt;code&amp;gt;doc/src/housekeeping/siv_status_bout.m4&amp;lt;/code&amp;gt;&lt;br /&gt;
all passed without error.&lt;br /&gt;
&lt;br /&gt;
== (#237) SUBADULT_ARRIVALS_LOG FirstTikiDate is declared NOT NULL contrary to its own documentation ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;db/schemas/housekeeping/tables/create/subadult_arrivals_log.m4&amp;lt;/code&amp;gt; declares&lt;br /&gt;
&amp;lt;code&amp;gt;firsttikidate DATE NOT NULL&amp;lt;/code&amp;gt;, but the owning documentation, added in the&lt;br /&gt;
same commit, says the column &amp;quot;may be NULL when there is no information on&lt;br /&gt;
the date of permanent addition.&amp;quot;  156 of 209 source rows (75%) have a NULL&lt;br /&gt;
&amp;lt;code&amp;gt;SA_first_tiki_date&amp;lt;/code&amp;gt;, each carrying an explanatory note describing why no&lt;br /&gt;
discrete tiki date applies.  A NULL &amp;lt;code&amp;gt;FirstTikiDate&amp;lt;/code&amp;gt; is this table&amp;#039;s normal,&lt;br /&gt;
permanent representation for most individuals, not missing data awaiting&lt;br /&gt;
entry.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT count(*) FILTER (WHERE sa_first_tiki_date IS NULL) AS null_date&lt;br /&gt;
     , count(*) FILTER (WHERE sa_first_tiki_date IS NOT NULL) AS dated&lt;br /&gt;
  FROM clean.subadult_arrivals_log;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns 156 NULL and 53 non-NULL rows.  &amp;lt;code&amp;gt;git show cc3640b&lt;br /&gt;
--stat&amp;lt;/code&amp;gt; shows the table, its documentation, its indexes, and its trigger&lt;br /&gt;
were all added in one commit, so this contradiction has existed since the&lt;br /&gt;
table was created and was never exercised by a loader.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: correct the schema defect&lt;br /&gt;
rather than exclude the 156 legitimately NULL-dated rows.  Remove &amp;lt;code&amp;gt;NOT&lt;br /&gt;
NULL&amp;lt;/code&amp;gt; from the &amp;lt;code&amp;gt;firsttikidate&amp;lt;/code&amp;gt; column in the owning create M4, regenerate&lt;br /&gt;
the owned table SQL artifact, and verify the corrected nullability in a&lt;br /&gt;
disposable or migrated database before loading.  Leave the unique &amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt;&lt;br /&gt;
index, the &amp;lt;code&amp;gt;FirstTikiDate&amp;lt;/code&amp;gt; index, the &amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt; foreign key and &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;&lt;br /&gt;
check, and the update trigger unchanged.  This is a target-schema&lt;br /&gt;
correction, not a clean-stage mutation, and excludes no source rows on its&lt;br /&gt;
own.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  &amp;lt;code&amp;gt;NOT NULL&amp;lt;/code&amp;gt; was removed from&lt;br /&gt;
&amp;lt;code&amp;gt;firsttikidate&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;db/schemas/housekeeping/tables/create/subadult_arrivals_log.m4&amp;lt;/code&amp;gt;; the owned&lt;br /&gt;
table SQL artifact was regenerated with &amp;lt;code&amp;gt;make db/schemas/createtables.sql&amp;lt;/code&amp;gt;&lt;br /&gt;
and shows &amp;lt;code&amp;gt;firsttikidate DATE&amp;lt;/code&amp;gt; with no other column or constraint changed.&lt;br /&gt;
&amp;lt;code&amp;gt;ALTER TABLE housekeeping.subadult_arrivals_log ALTER COLUMN firsttikidate&lt;br /&gt;
DROP NOT NULL&amp;lt;/code&amp;gt; was applied to the local PostgreSQL 18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; disposable&lt;br /&gt;
database, and &amp;lt;code&amp;gt;information_schema.columns&amp;lt;/code&amp;gt; confirmed &amp;lt;code&amp;gt;firsttikidate&amp;lt;/code&amp;gt; is&lt;br /&gt;
nullable while &amp;lt;code&amp;gt;id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;animid&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;notes&amp;lt;/code&amp;gt; remain &amp;lt;code&amp;gt;NOT NULL&amp;lt;/code&amp;gt; and the unique&lt;br /&gt;
&amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt; index, &amp;lt;code&amp;gt;FirstTikiDate&amp;lt;/code&amp;gt; index, &amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt; foreign key and &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;&lt;br /&gt;
check, and update trigger are unchanged.&lt;br /&gt;
&lt;br /&gt;
== * (#238) SUBADULT_ARRIVALS_LOG CT&amp;#039;s FirstTikiDate follows CT&amp;#039;s own departure date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;CT&amp;lt;/code&amp;gt;&amp;#039;s parsed &amp;lt;code&amp;gt;FirstTikiDate&amp;lt;/code&amp;gt; (1986-03-01) is after &amp;lt;code&amp;gt;CT&amp;lt;/code&amp;gt;&amp;#039;s own&lt;br /&gt;
&amp;lt;code&amp;gt;BIOGRAPHY_DATA.DepartDate&amp;lt;/code&amp;gt; (1985-09-30).  The target has no constraint&lt;br /&gt;
tying &amp;lt;code&amp;gt;FirstTikiDate&amp;lt;/code&amp;gt; to biography dates, so this row would load without&lt;br /&gt;
error, but a tiki date recorded after an individual&amp;#039;s departure is not&lt;br /&gt;
trustworthy evidence of when that individual was added to the tiki sheets.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT source.*, biography.departdate&lt;br /&gt;
  FROM clean.subadult_arrivals_log AS source&lt;br /&gt;
  JOIN sokwedb.biography_data AS biography&lt;br /&gt;
    ON biography.animid = source.sa_b_animid&lt;br /&gt;
  WHERE to_date(source.sa_first_tiki_date, &amp;#039;MM/DD/YY&amp;#039;) &amp;gt; biography.departdate;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns exactly one row: &amp;lt;code&amp;gt;CT&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;03/01/86&amp;lt;/code&amp;gt;, with note &amp;quot;Btwn&lt;br /&gt;
DOB and CA death, arrivals same as CA; Elo confirmed that always accounted&lt;br /&gt;
for on tikis afterwards.&amp;quot;  No other row falls outside its individual&amp;#039;s&lt;br /&gt;
birth/entry/departure window.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: exclude this exact row rather&lt;br /&gt;
than preserve it as a documented semantic anomaly.  Match on the exact&lt;br /&gt;
source tuple (&amp;lt;code&amp;gt;sa_b_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;sa_first_tiki_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;sa_notes&amp;lt;/code&amp;gt;) with&lt;br /&gt;
NULL-safe equality; do not exclude by &amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt; alone or use a broad&lt;br /&gt;
date-comparison predicate as the operative exclusion, and do not infer a&lt;br /&gt;
corrected date.  This is the only approved exclusion; it leaves 208&lt;br /&gt;
eligible rows.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/subadult_arrivals_log_projection.sql&amp;lt;/code&amp;gt; encodes the exact&lt;br /&gt;
one-row &amp;lt;code&amp;gt;CT&amp;lt;/code&amp;gt; tuple.  Read-only sanity, run against the local PostgreSQL&lt;br /&gt;
18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; database, reproduced this exact tuple with bidirectional&lt;br /&gt;
&amp;lt;code&amp;gt;EXCEPT ALL&amp;lt;/code&amp;gt; and confirmed no other eligible row falls outside its&lt;br /&gt;
individual&amp;#039;s birth/entry/departure window.&lt;br /&gt;
&lt;br /&gt;
== (#239) SUBADULT_ARRIVALS_LOG absent notes cannot satisfy the target ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The source contains 6 NULL &amp;lt;code&amp;gt;SA_notes&amp;lt;/code&amp;gt; values, all outside the Problem #238&lt;br /&gt;
exclusion.  The target &amp;lt;code&amp;gt;Notes&amp;lt;/code&amp;gt; is NOT NULL but explicitly permits the empty&lt;br /&gt;
string.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT count(*)&lt;br /&gt;
  FROM clean.subadult_arrivals_log&lt;br /&gt;
  WHERE sa_notes IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns 6.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: map an eligible NULL note to&lt;br /&gt;
the empty string with &amp;lt;code&amp;gt;COALESCE(source.sa_notes, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt;.  Preserve every&lt;br /&gt;
non-NULL note exactly, including two rows (&amp;lt;code&amp;gt;FAD&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;KEA&amp;lt;/code&amp;gt;) that carry leading&lt;br /&gt;
or trailing spaces around substantive text.  Sanity and validation must&lt;br /&gt;
prove that this is the only notes transformation and that no projected&lt;br /&gt;
note contains only spaces.  This mapping, together with Problem #237&amp;#039;s&lt;br /&gt;
schema fix and Problem #238&amp;#039;s exclusion, leaves 208 rows for conversion:&lt;br /&gt;
209 source rows minus the one Problem #238 exclusion.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.&lt;br /&gt;
&amp;lt;code&amp;gt;COALESCE(source.sa_notes, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt; is the projection&amp;#039;s only notes&lt;br /&gt;
transformation. Read-only sanity and validation against the local&lt;br /&gt;
PostgreSQL 18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; database confirmed 6 loaded empty notes came&lt;br /&gt;
only from NULL source notes, all 202 non-NULL notes are preserved&lt;br /&gt;
byte-for-byte (including the two rows with leading/trailing spaces), and&lt;br /&gt;
no loaded note contains only spaces.&lt;br /&gt;
&lt;br /&gt;
The direct set-based loader (&amp;lt;code&amp;gt;conversion/load_subadult_arrivals_log.sql&amp;lt;/code&amp;gt;)&lt;br /&gt;
inserted exactly 208 rows: 156 NULL and 52 non-NULL &amp;lt;code&amp;gt;FirstTikiDate&amp;lt;/code&amp;gt;&lt;br /&gt;
values (extent 1976-05-01 through 2013-09-22, unchanged by the &amp;lt;code&amp;gt;CT&amp;lt;/code&amp;gt;&lt;br /&gt;
exclusion), and 6 empty/202 nonempty notes. Every non-NULL &amp;lt;code&amp;gt;FirstTikiDate&amp;lt;/code&amp;gt;&lt;br /&gt;
was proven to round-trip through &amp;lt;code&amp;gt;to_char(..., &amp;#039;MM/DD/YY&amp;#039;)&amp;lt;/code&amp;gt; back to its&lt;br /&gt;
exact source text. Bidirectional &amp;lt;code&amp;gt;EXCEPT ALL&amp;lt;/code&amp;gt; between the canonical&lt;br /&gt;
projection and the loaded rows returned no differences in either&lt;br /&gt;
direction. An intentional post-insert failure (&amp;lt;code&amp;gt;SELECT 1/0&amp;lt;/code&amp;gt;) rolled the&lt;br /&gt;
single transaction back to zero target rows. An authorized &amp;lt;code&amp;gt;TRUNCATE&amp;lt;/code&amp;gt; of&lt;br /&gt;
only the disposable target table followed by an unchanged&lt;br /&gt;
sanity/load/validation sequence reproduced the identical sorted&lt;br /&gt;
four-column payload digest, &amp;lt;code&amp;gt;6eac68b10ebb802e3c78b9109e28d3d7&amp;lt;/code&amp;gt;, on both&lt;br /&gt;
runs. &amp;lt;code&amp;gt;git diff --check&amp;lt;/code&amp;gt;, a local-only Make dry run of the full&lt;br /&gt;
&amp;lt;code&amp;gt;load_data&amp;lt;/code&amp;gt; sequence, and a direct &amp;lt;code&amp;gt;m4&amp;lt;/code&amp;gt; and repository doc-build render&lt;br /&gt;
of &amp;lt;code&amp;gt;doc/src/housekeeping/subadult_arrivals_log.m4&amp;lt;/code&amp;gt; all passed without&lt;br /&gt;
error.&lt;br /&gt;
&lt;br /&gt;
== (#240) TLK_BRECORD_NOTES_CODES two Abbreviation values carry a trailing space ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 94&amp;#039;s &amp;lt;code&amp;gt;&amp;quot;Abbreviation&amp;quot;&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;l &amp;lt;/code&amp;gt; (trailing space) with &amp;lt;code&amp;gt;&amp;quot;Meaning&amp;quot;&amp;lt;/code&amp;gt; of&lt;br /&gt;
&amp;lt;code&amp;gt;laugh&amp;lt;/code&amp;gt;; &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 186&amp;#039;s &amp;lt;code&amp;gt;&amp;quot;Abbreviation&amp;quot;&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;th &amp;lt;/code&amp;gt; (trailing space) with&lt;br /&gt;
&amp;lt;code&amp;gt;&amp;quot;Meaning&amp;quot;&amp;lt;/code&amp;gt; of &amp;lt;code&amp;gt;threaten&amp;lt;/code&amp;gt;.  Both violate the target&amp;#039;s&lt;br /&gt;
&amp;lt;code&amp;gt;trimmedofspaces_check&amp;lt;/code&amp;gt;.  Trimming either value makes it collide with an&lt;br /&gt;
existing unpadded row (&amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 93 &amp;lt;code&amp;gt;l&amp;lt;/code&amp;gt; -&amp;gt; &amp;lt;code&amp;gt;locomotion/travel&amp;lt;/code&amp;gt;; &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 185&lt;br /&gt;
&amp;lt;code&amp;gt;th&amp;lt;/code&amp;gt; -&amp;gt; &amp;lt;code&amp;gt;throw&amp;lt;/code&amp;gt;), folding directly into the Problem #241 duplicate-&lt;br /&gt;
abbreviation resolution.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;ID&amp;quot;, &amp;quot;Abbreviation&amp;quot;, &amp;quot;Meaning&amp;quot;&lt;br /&gt;
  FROM clean.tlk_brecord_notes_codes&lt;br /&gt;
  WHERE &amp;quot;Abbreviation&amp;quot; IS DISTINCT FROM BTRIM(&amp;quot;Abbreviation&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns exactly two rows: &amp;lt;code&amp;gt;94, &amp;quot;l &amp;quot;, laugh&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;186,&lt;br /&gt;
&amp;quot;th &amp;quot;, threaten&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19, as an explicit exception to&lt;br /&gt;
this repository&amp;#039;s usual prohibition on clean-stage mutation: trim&lt;br /&gt;
&amp;lt;code&amp;gt;&amp;quot;Abbreviation&amp;quot;&amp;lt;/code&amp;gt; in &amp;lt;code&amp;gt;conversion/clean.sql&amp;lt;/code&amp;gt;, in its own reviewable&lt;br /&gt;
statement, rather than in the loader or projection.  An initial&lt;br /&gt;
implementation matched the &amp;lt;code&amp;gt;UPDATE&amp;lt;/code&amp;gt; by exact &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot; IN (94, 186)&amp;lt;/code&amp;gt;; on&lt;br /&gt;
review the investigator preferred applying &amp;lt;code&amp;gt;BTRIM&amp;lt;/code&amp;gt; unconditionally, as a&lt;br /&gt;
general clean-stage normalization, since it produces the identical result&lt;br /&gt;
here (no other row is padded) and reads more simply as ordinary cleanup&lt;br /&gt;
rather than a two-row special case.  The dated source affects exactly the&lt;br /&gt;
same two rows either way, &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 94 and &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 186.  After this fix, &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt;&lt;br /&gt;
94&amp;#039;s abbreviation is &amp;lt;code&amp;gt;l&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 186&amp;#039;s is &amp;lt;code&amp;gt;th&amp;lt;/code&amp;gt;, each now a member of an&lt;br /&gt;
ordinary Problem #241 duplicate group.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  Added an unconditional&lt;br /&gt;
&amp;lt;code&amp;gt;UPDATE ... SET &amp;quot;Abbreviation&amp;quot; = BTRIM(&amp;quot;Abbreviation&amp;quot;)&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/clean.sql&amp;lt;/code&amp;gt;, independently reviewable from the loader, with a&lt;br /&gt;
comment documenting that the dated source&amp;#039;s only two affected rows are&lt;br /&gt;
&amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 94 and &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 186.  Read-only re-checks against the local&lt;br /&gt;
PostgreSQL 18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; database confirmed the two exact pre-fix rows&lt;br /&gt;
before the trim, that applying the statement leaves zero padded rows&lt;br /&gt;
afterward, and that the two documented rows&amp;#039; post-fix unpadded values&lt;br /&gt;
(&amp;lt;code&amp;gt;l&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;th&amp;lt;/code&amp;gt;) are unchanged from what the narrowly-targeted form would have&lt;br /&gt;
produced.&lt;br /&gt;
&lt;br /&gt;
== * (#241) TLK_BRECORD_NOTES_CODES duplicate abbreviations after trimming ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After Problem #240&amp;#039;s fix, 26 case-insensitive abbreviations occur more than&lt;br /&gt;
once, with genuinely different &amp;lt;code&amp;gt;Meaning&amp;lt;/code&amp;gt; text recorded for the same&lt;br /&gt;
abbreviation depending on note-taking context (for example &amp;lt;code&amp;gt;r&amp;lt;/code&amp;gt; means&lt;br /&gt;
&amp;lt;code&amp;gt;rest&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;reassurance&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;reunion&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;resume&amp;lt;/code&amp;gt; across four separate rows).&lt;br /&gt;
The target permits only one row per case-insensitive abbreviation via a&lt;br /&gt;
unique index on &amp;lt;code&amp;gt;lower(normalize(Abbreviation))&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT lower(normalize(BTRIM(&amp;quot;Abbreviation&amp;quot;))) AS norm_abbr, count(*)&lt;br /&gt;
     , array_agg(&amp;quot;ID&amp;quot; ORDER BY &amp;quot;ID&amp;quot;) AS ids&lt;br /&gt;
     , array_agg(&amp;quot;Meaning&amp;quot; ORDER BY &amp;quot;ID&amp;quot;) AS meanings&lt;br /&gt;
  FROM clean.tlk_brecord_notes_codes&lt;br /&gt;
  GROUP BY 1&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns exactly 26 groups covering 55 rows: 24 groups of 2&lt;br /&gt;
rows, one group of 3 (&amp;lt;code&amp;gt;ds&amp;lt;/code&amp;gt;), and one group of 4 (&amp;lt;code&amp;gt;r&amp;lt;/code&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Superseded on 2026-09-19.  An initial approval picked the lowest source&lt;br /&gt;
&amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; in each group to survive.  On review the investigator rejected that&lt;br /&gt;
tie-break: choosing a survivor&amp;#039;s &amp;lt;code&amp;gt;Meaning&amp;lt;/code&amp;gt; over its siblings&amp;#039; is itself an&lt;br /&gt;
uncorroborated judgment call about which recorded meaning is &amp;quot;correct&amp;quot;,&lt;br /&gt;
which is exactly what a targeted, documented exclusion is meant to avoid.&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19 (final): exclude every row in&lt;br /&gt;
each of the 26 case-insensitive duplicate-abbreviation groups, not only the&lt;br /&gt;
non-survivors.  None of the 26 ambiguous abbreviations is loaded under any&lt;br /&gt;
&amp;lt;code&amp;gt;Meaning&amp;lt;/code&amp;gt;; only the 155 rows whose abbreviation is unique in the source&lt;br /&gt;
(after Problem #240&amp;#039;s trim) are eligible.  This drops all 55 grouped rows,&lt;br /&gt;
not 29.  The excluded rows, and their meanings, remain visible, unmodified,&lt;br /&gt;
in &amp;lt;code&amp;gt;clean.tlk_brecord_notes_codes&amp;lt;/code&amp;gt;.  Do not infer, rank, or merge meanings;&lt;br /&gt;
do not pick a survivor by &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; or any other criterion.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  &amp;lt;code&amp;gt;conversion/tlk_brecord_notes_codes_projection.sql&amp;lt;/code&amp;gt;&lt;br /&gt;
excludes every row whose case-insensitive abbreviation occurs more than&lt;br /&gt;
once in &amp;lt;code&amp;gt;clean.tlk_brecord_notes_codes&amp;lt;/code&amp;gt; (after the Problem #240 trim),&lt;br /&gt;
proved against an independently-derived 55-row, 26-abbreviation exclusion&lt;br /&gt;
set.  Read-only sanity (&amp;lt;code&amp;gt;conversion/load_tlk_brecord_notes_codes_sanity.sql&amp;lt;/code&amp;gt;)&lt;br /&gt;
against the local PostgreSQL 18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; database confirmed the&lt;br /&gt;
exclusion set matches exactly and that the canonical projection is exactly&lt;br /&gt;
155 rows with 155 distinct case-insensitive abbreviations, none of them one&lt;br /&gt;
of the 26 excluded abbreviations.&lt;br /&gt;
&lt;br /&gt;
The direct set-based loader (&amp;lt;code&amp;gt;conversion/load_tlk_brecord_notes_codes.sql&amp;lt;/code&amp;gt;)&lt;br /&gt;
inserted exactly 155 rows.  Bidirectional &amp;lt;code&amp;gt;EXCEPT ALL&amp;lt;/code&amp;gt; between the&lt;br /&gt;
canonical projection and the loaded rows&lt;br /&gt;
(&amp;lt;code&amp;gt;conversion/load_tlk_brecord_notes_codes_validation.sql&amp;lt;/code&amp;gt;) returned no&lt;br /&gt;
differences in either direction, every loaded abbreviation was proven&lt;br /&gt;
unique under &amp;lt;code&amp;gt;lower(normalize(...))&amp;lt;/code&amp;gt; and to have exactly one matching&lt;br /&gt;
source row, every loaded &amp;lt;code&amp;gt;Meaning&amp;lt;/code&amp;gt; was proven to match that single source&lt;br /&gt;
row exactly, and none of the 26 excluded abbreviations was proven present&lt;br /&gt;
in the loaded rows.  An intentional post-insert failure (&amp;lt;code&amp;gt;SELECT 1/0&amp;lt;/code&amp;gt;)&lt;br /&gt;
rolled the single transaction back to zero target rows.  An authorized&lt;br /&gt;
&amp;lt;code&amp;gt;TRUNCATE&amp;lt;/code&amp;gt; of only the disposable target table followed by an unchanged&lt;br /&gt;
sanity/load/validation sequence, run twice, reproduced the identical&lt;br /&gt;
sorted two-column payload digest, &amp;lt;code&amp;gt;27da98a74e6378218418652dda914c95&amp;lt;/code&amp;gt;, on&lt;br /&gt;
both runs.  &amp;lt;code&amp;gt;git diff --check&amp;lt;/code&amp;gt;, a&lt;br /&gt;
local-only Make dry run of the full &amp;lt;code&amp;gt;load_data&amp;lt;/code&amp;gt; sequence confirming correct&lt;br /&gt;
ordering, a direct Make-driven execution of the new&lt;br /&gt;
&amp;lt;code&amp;gt;load_tlk_brecord_notes_codes_sanity&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;load_tlk_brecord_notes_codes&amp;lt;/code&amp;gt;/&lt;br /&gt;
&amp;lt;code&amp;gt;load_tlk_brecord_notes_codes_validation&amp;lt;/code&amp;gt; targets, and a direct &amp;lt;code&amp;gt;m4&amp;lt;/code&amp;gt; render&lt;br /&gt;
of &amp;lt;code&amp;gt;doc/src/housekeeping/tlk_brecord_notes_codes.m4&amp;lt;/code&amp;gt; all passed without&lt;br /&gt;
error.&lt;br /&gt;
&lt;br /&gt;
== * (#242) FOOD_VARIATIONS eligible rows lack a LocalName ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Problem #161 deferred all of &amp;lt;code&amp;gt;clean.food_variations_lookup&amp;lt;/code&amp;gt; to a future&lt;br /&gt;
conversion because of missing and unresolved &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt; values.  Picking&lt;br /&gt;
this conversion back up: the target &amp;lt;code&amp;gt;housekeeping.food_variations.LocalName&amp;lt;/code&amp;gt;&lt;br /&gt;
is &amp;lt;code&amp;gt;NOT NULL&amp;lt;/code&amp;gt;, but 224 of the source&amp;#039;s 1,575 rows have a NULL&lt;br /&gt;
&amp;lt;code&amp;gt;fvl_fl_local_food_name&amp;lt;/code&amp;gt;.  40 of those 224 rows have a&lt;br /&gt;
&amp;lt;code&amp;gt;fvl_food_spelling_variant&amp;lt;/code&amp;gt; that itself exactly matches an existing&lt;br /&gt;
&amp;lt;code&amp;gt;codes.food_names&amp;lt;/code&amp;gt; entry (for example &amp;lt;code&amp;gt;DONGO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;KIFUMUFUMU&amp;lt;/code&amp;gt;), but the&lt;br /&gt;
other 184 do not, so there is no single reliable substitute value across&lt;br /&gt;
all 224 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT count(*) AS null_localname_rows&lt;br /&gt;
     , count(*) FILTER (&lt;br /&gt;
         WHERE EXISTS (&lt;br /&gt;
           SELECT 1 FROM codes.food_names AS fn&lt;br /&gt;
           WHERE lower(normalize(fn.name)) =&lt;br /&gt;
                   lower(normalize(fvl.fvl_food_spelling_variant))))&lt;br /&gt;
         AS variant_matches_food_names&lt;br /&gt;
  FROM clean.food_variations_lookup AS fvl&lt;br /&gt;
  WHERE fvl.fvl_fl_local_food_name IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-19 snapshot returns 224 and 40.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: exclude all 224 rows with a&lt;br /&gt;
NULL &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt;, including the 40 whose &amp;lt;code&amp;gt;Variant&amp;lt;/code&amp;gt; happens to match a&lt;br /&gt;
&amp;lt;code&amp;gt;codes.food_names&amp;lt;/code&amp;gt; entry.  Do not infer a &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt; from the &amp;lt;code&amp;gt;Variant&amp;lt;/code&amp;gt;,&lt;br /&gt;
even for those 40 rows.  This is a straightforward domain predicate&lt;br /&gt;
(&amp;lt;code&amp;gt;fvl_fl_local_food_name IS NOT NULL&amp;lt;/code&amp;gt;), not an exact-tuple exclusion list,&lt;br /&gt;
since excluding by this condition cannot ambiguously over- or&lt;br /&gt;
under-exclude the way a duplicate-abbreviation tie-break could.  This&lt;br /&gt;
leaves at most 1,351 eligible rows before any other approved exclusion.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-20.  &amp;lt;code&amp;gt;conversion/food_variations_projection.sql&amp;lt;/code&amp;gt;&lt;br /&gt;
excludes every row via the domain predicate &amp;lt;code&amp;gt;fvl_fl_local_food_name IS&lt;br /&gt;
NOT NULL&amp;lt;/code&amp;gt;.  Read-only sanity (&amp;lt;code&amp;gt;conversion/load_food_variations_sanity.sql&amp;lt;/code&amp;gt;)&lt;br /&gt;
against the local PostgreSQL 18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; database confirmed the exact&lt;br /&gt;
224-row NULL-&amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt; count and that the canonical projection is&lt;br /&gt;
exactly 1,351 rows with 1,351 distinct case-insensitive &amp;lt;code&amp;gt;Variant&amp;lt;/code&amp;gt; values.&lt;br /&gt;
The direct set-based loader (&amp;lt;code&amp;gt;conversion/load_food_variations.sql&amp;lt;/code&amp;gt;)&lt;br /&gt;
inserted exactly 1,351 rows.  Bidirectional &amp;lt;code&amp;gt;EXCEPT ALL&amp;lt;/code&amp;gt; between the&lt;br /&gt;
canonical projection and the loaded rows&lt;br /&gt;
(&amp;lt;code&amp;gt;conversion/load_food_variations_validation.sql&amp;lt;/code&amp;gt;) returned no&lt;br /&gt;
differences in either direction, every loaded &amp;lt;code&amp;gt;Variant&amp;lt;/code&amp;gt; was proven unique&lt;br /&gt;
under &amp;lt;code&amp;gt;lower(normalize(...))&amp;lt;/code&amp;gt;, every loaded &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt; was proven to&lt;br /&gt;
match its source row exactly, and no loaded row was proven to have a NULL&lt;br /&gt;
source &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt;.  An intentional post-insert failure (&amp;lt;code&amp;gt;SELECT 1/0&amp;lt;/code&amp;gt;)&lt;br /&gt;
rolled the single transaction back to zero target rows.  An authorized&lt;br /&gt;
&amp;lt;code&amp;gt;TRUNCATE&amp;lt;/code&amp;gt; of only the disposable target table followed by an unchanged&lt;br /&gt;
sanity/load/validation sequence, run twice, reproduced the identical&lt;br /&gt;
sorted two-column payload digest, &amp;lt;code&amp;gt;c510323457dcb801412441eda0ceb09c&amp;lt;/code&amp;gt;, on&lt;br /&gt;
both runs.  &amp;lt;code&amp;gt;git diff --check&amp;lt;/code&amp;gt;, a local-only Make dry run of the full&lt;br /&gt;
&amp;lt;code&amp;gt;load_data&amp;lt;/code&amp;gt; sequence confirming correct ordering, a direct Make-driven&lt;br /&gt;
execution of the new &amp;lt;code&amp;gt;load_food_variations_sanity&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;load_food_variations&amp;lt;/code&amp;gt;/&lt;br /&gt;
&amp;lt;code&amp;gt;load_food_variations_validation&amp;lt;/code&amp;gt; targets, and a direct &amp;lt;code&amp;gt;m4&amp;lt;/code&amp;gt; render of&lt;br /&gt;
&amp;lt;code&amp;gt;doc/src/housekeeping/food_variations.m4&amp;lt;/code&amp;gt; all passed without error.&lt;br /&gt;
&lt;br /&gt;
== (#243) FOOD_VARIATIONS orphaned LocalName values are not cross-checked ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Of the 1,351 rows surviving Problem #242, 38 have a &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt; that does&lt;br /&gt;
not match any &amp;lt;code&amp;gt;clean.food_lookup&amp;lt;/code&amp;gt; entry case-insensitively, and a&lt;br /&gt;
stricter 79 do not match the smaller, already-loaded production&lt;br /&gt;
&amp;lt;code&amp;gt;codes.food_names&amp;lt;/code&amp;gt; table.  &amp;lt;code&amp;gt;housekeeping.food_variations&amp;lt;/code&amp;gt; declares no&lt;br /&gt;
foreign key to either table, so nothing in the schema forces a&lt;br /&gt;
resolution; several of the mismatches form typo clusters pointing at a&lt;br /&gt;
name that is not itself canonical either (for example eight variants —&lt;br /&gt;
&amp;lt;code&amp;gt;MBOBOGOLO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MTOMBOGOLO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MUTBOGORO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MUTOBOGALO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MUTOBOGO&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;MUTOBOLO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MUTOBORGORO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MUTOLGURO&amp;lt;/code&amp;gt; — all record &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;MTOBOGOLO&amp;lt;/code&amp;gt;, which appears in neither reference table).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT count(*) FILTER (&lt;br /&gt;
         WHERE NOT EXISTS (&lt;br /&gt;
           SELECT 1 FROM clean.food_lookup AS fl&lt;br /&gt;
           WHERE lower(fl.fl_local_food_name) = lower(fvl.fvl_fl_local_food_name)))&lt;br /&gt;
         AS orphan_vs_food_lookup&lt;br /&gt;
     , count(*) FILTER (&lt;br /&gt;
         WHERE NOT EXISTS (&lt;br /&gt;
           SELECT 1 FROM codes.food_names AS fn&lt;br /&gt;
           WHERE lower(normalize(fn.name)) =&lt;br /&gt;
                   lower(normalize(fvl.fvl_fl_local_food_name))))&lt;br /&gt;
         AS orphan_vs_food_names&lt;br /&gt;
  FROM clean.food_variations_lookup AS fvl&lt;br /&gt;
  WHERE fvl.fvl_fl_local_food_name IS NOT NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-20 snapshot returns 38 and 79, unchanged since 2026-09-19.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-20: do not cross-check &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt;&lt;br /&gt;
against &amp;lt;code&amp;gt;clean.food_lookup&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;codes.food_names&amp;lt;/code&amp;gt;, or any other table for&lt;br /&gt;
this conversion.  Load every eligible row&amp;#039;s &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt; as free text&lt;br /&gt;
exactly as recorded, matching, mismatching, or orphaned alike.  Sanity may&lt;br /&gt;
report the 38/79 counts as informational profiling, but neither figure&lt;br /&gt;
excludes any row.  This mapping excludes no additional rows beyond&lt;br /&gt;
Problem #242&amp;#039;s 224.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-20.  The canonical projection applies&lt;br /&gt;
no cross-check against &amp;lt;code&amp;gt;clean.food_lookup&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;codes.food_names&amp;lt;/code&amp;gt;; sanity&lt;br /&gt;
reports the 38/79 mismatch counts as &amp;lt;code&amp;gt;RAISE NOTICE&amp;lt;/code&amp;gt; profiling only.&lt;br /&gt;
Validation confirmed every loaded &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt; matches its source row&lt;br /&gt;
exactly regardless of whether it matches either reference table, and the&lt;br /&gt;
1,351-row loaded count matches &amp;lt;code&amp;gt;1,575 source - 224 Problem #242&lt;br /&gt;
exclusions&amp;lt;/code&amp;gt; exactly, confirming no additional row was excluded under this&lt;br /&gt;
resolution.&lt;br /&gt;
&lt;br /&gt;
== (#244) REPRO_STATES cannot represent adolescent and not-seen days ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;clean.female_reproductive_states.repro_state&amp;lt;/code&amp;gt; contains &amp;lt;code&amp;gt;A&amp;lt;/code&amp;gt; for adolescent&lt;br /&gt;
and SQL NULL for not seen, but production &amp;lt;code&amp;gt;REPRO_STATES.State&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NOT NULL&amp;lt;/code&amp;gt;&lt;br /&gt;
and permits only cycling (&amp;lt;code&amp;gt;C&amp;lt;/code&amp;gt;), pregnant (&amp;lt;code&amp;gt;P&amp;lt;/code&amp;gt;), and lactating (&amp;lt;code&amp;gt;L&amp;lt;/code&amp;gt;).  The&lt;br /&gt;
dated 2026-09-19 source has 7,582 &amp;lt;code&amp;gt;A&amp;lt;/code&amp;gt; rows and 218,370 NULL rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT repro_state, count(*)&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  GROUP BY repro_state&lt;br /&gt;
  ORDER BY repro_state NULLS LAST;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: adolescent &amp;lt;code&amp;gt;A&amp;lt;/code&amp;gt;&lt;br /&gt;
was added to the target state domain and &amp;lt;code&amp;gt;REPRO_STATES.State&amp;lt;/code&amp;gt; was made&lt;br /&gt;
nullable.  Map source &amp;lt;code&amp;gt;A&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;C&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;P&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;LA&amp;lt;/code&amp;gt; to target &amp;lt;code&amp;gt;A&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;C&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;P&amp;lt;/code&amp;gt;, and&lt;br /&gt;
&amp;lt;code&amp;gt;L&amp;lt;/code&amp;gt;; preserve source NULL as target NULL.  Do not carry a previous state&lt;br /&gt;
into a not-seen day.  The generated SQL and documentation were regenerated;&lt;br /&gt;
catalog inspection and a rollback-only NULL-State insert verified the change.&lt;br /&gt;
&lt;br /&gt;
== (#245) REPRO_STATES requires sparse support attributes ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;Source&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;EstrusDayCert&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;LactEndCert&amp;lt;/code&amp;gt; are &amp;lt;code&amp;gt;NOT NULL&amp;lt;/code&amp;gt;, but&lt;br /&gt;
the dated source contains 646,257, 626,192, and 646,602 NULL values&lt;br /&gt;
respectively.  The 30 rows having an EstrusDay but NULL EstrusDayCert are&lt;br /&gt;
valid under the approved sparse representation.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT count(*) FILTER (WHERE state_change_dec_source IS NULL) AS null_source&lt;br /&gt;
     , count(*) FILTER (WHERE day_of_est_cert IS NULL) AS null_ed_cert&lt;br /&gt;
     , count(*) FILTER (WHERE lactat_end_cert IS NULL) AS null_le_cert&lt;br /&gt;
     , count(*) FILTER (&lt;br /&gt;
         WHERE day_of_estr IS NOT NULL AND day_of_est_cert IS NULL)&lt;br /&gt;
         AS estrus_day_without_cert&lt;br /&gt;
  FROM clean.female_reproductive_states;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: the three&lt;br /&gt;
production columns were made nullable.  Preserve source NULL as SQL NULL;&lt;br /&gt;
do not carry values forward and do not convert NULL to empty text.  The&lt;br /&gt;
support tables reject empty keys, and an empty key would invent a category&lt;br /&gt;
rather than represent missing information.  Catalog inspection and a&lt;br /&gt;
rollback-only sparse-row insert verified the change.&lt;br /&gt;
&lt;br /&gt;
== (#246) REPRO_STATES support tables are empty ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The four required support tables are empty.  The source has twelve distinct&lt;br /&gt;
trimmed change-source labels, ED certainty values &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt; plus the&lt;br /&gt;
Problem #247 typo, LE certainty values &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt;, and three parity&lt;br /&gt;
values.  The Access dump contains no lookup tables defining them.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT upper(btrim(state_change_dec_source)) AS code, count(*)&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  WHERE state_change_dec_source IS NOT NULL&lt;br /&gt;
  GROUP BY 1&lt;br /&gt;
  ORDER BY 1;&lt;br /&gt;
&lt;br /&gt;
SELECT parity, count(*)&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  GROUP BY parity&lt;br /&gt;
  ORDER BY parity;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: load uppercase&lt;br /&gt;
parity keys &amp;lt;code&amp;gt;IMMATURE&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;NULLIPAROUS&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;PAROUS&amp;lt;/code&amp;gt;.  Normalize populated&lt;br /&gt;
change-source keys with &amp;lt;code&amp;gt;upper(btrim(...))&amp;lt;/code&amp;gt;, retaining the twelve semantic&lt;br /&gt;
labels and expanding their state prefixes in Description; the exact&lt;br /&gt;
key/description pairs are recorded in &amp;lt;code&amp;gt;repro_states_loader_handoff.md&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Load ED keys &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt; with the approved descriptions based on the&lt;br /&gt;
length of the unobserved gap following the last full swelling.  Load LE keys&lt;br /&gt;
&amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt; with the approved minimum/maximum swelling, 14-day, and&lt;br /&gt;
first-cycle criteria.  ED code &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt; has no support row; see Problem #248, which&lt;br /&gt;
now corrects &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt; in the clean schema rather than excluding it, so no&lt;br /&gt;
row ever carries an unsupported code.  No support row represents NULL; see&lt;br /&gt;
Problem #245.  Sanity verifies bidirectional parity of every support key and&lt;br /&gt;
description; the clean retry loaded exactly 12 change sources, four ED&lt;br /&gt;
certainties, four LE certainties, and three parities.&lt;br /&gt;
&lt;br /&gt;
== (#247) FEMALE_REPRODUCTIVE_STATES has estrus certainty 22 ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
One row, SIF on 2004-03-21, has &amp;lt;code&amp;gt;day_of_est_cert = &amp;#039;22&amp;#039;&amp;lt;/code&amp;gt;.  The investigator&lt;br /&gt;
confirmed that this is a typo for &amp;lt;code&amp;gt;2&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  WHERE day_of_est_cert = &amp;#039;22&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: change &amp;lt;code&amp;gt;22&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;2&amp;lt;/code&amp;gt; in the clean schema.  The focused clean-stage execution changed exactly&lt;br /&gt;
one row; sanity verifies that &amp;lt;code&amp;gt;22&amp;lt;/code&amp;gt; is absent and the corrected &amp;lt;code&amp;gt;2&amp;lt;/code&amp;gt; count is&lt;br /&gt;
8,251.&lt;br /&gt;
&lt;br /&gt;
== (#248) FEMALE_REPRODUCTIVE_STATES estrus certainty 5 is unsupported ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The dated source has 30 rows with &amp;lt;code&amp;gt;day_of_est_cert = &amp;#039;5&amp;#039;&amp;lt;/code&amp;gt;.  The investigator&lt;br /&gt;
provided meanings for ED certainty codes &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt;; code &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt; has no&lt;br /&gt;
supported meaning.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  WHERE day_of_est_cert = &amp;#039;5&amp;#039;&lt;br /&gt;
  ORDER BY id, date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: temporarily&lt;br /&gt;
exclude exactly these 30 rows from the reproductive-state projection.  Code&lt;br /&gt;
&amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt; is absent from &amp;lt;code&amp;gt;ED_CERTAINTIES&amp;lt;/code&amp;gt;.  Sanity verifies both the 30-row class and&lt;br /&gt;
that it overlaps none of Problems #251 or #254.&lt;br /&gt;
&lt;br /&gt;
Amended by the investigator on 2026-09-24: convert &amp;lt;code&amp;gt;day_of_est_cert = &amp;#039;5&amp;#039;&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt; in the clean schema instead of excluding the 30 rows.  No &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt; value is&lt;br /&gt;
retained, no row is excluded, and no support row is added for &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt;; the&lt;br /&gt;
correction is applied alongside the Problem #247 typo correction in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/clean.sql&amp;lt;/code&amp;gt;.  Sanity now asserts an exact 2,325-row &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt; count and&lt;br /&gt;
rejects any remaining &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt; value as a corrected-domain violation.  The&lt;br /&gt;
approved projection grew from 649,495 to 649,525 rows, and the pre-expansion&lt;br /&gt;
eligible/exclusion counts changed accordingly; the refreshed counts are&lt;br /&gt;
verified in &amp;lt;code&amp;gt;conversion/repro_states_profile.sql&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/load_repro_states_sanity.sql&amp;lt;/code&amp;gt;, and&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/load_repro_states_validation.sql&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#249) FEMALE_REPRODUCTIVE_STATES swelling bounds have no target ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Source columns &amp;lt;code&amp;gt;min_sw&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;max_sw&amp;lt;/code&amp;gt; have no &amp;lt;code&amp;gt;REPRO_STATES&amp;lt;/code&amp;gt; destination.  The&lt;br /&gt;
dated source has 160,268 rows where both are populated and 486,621 where both&lt;br /&gt;
are NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT min_sw, max_sw, count(*)&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  GROUP BY min_sw, max_sw&lt;br /&gt;
  ORDER BY min_sw NULLS LAST, max_sw NULLS LAST;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: do not map or&lt;br /&gt;
convert these fields because they are generated by &amp;lt;code&amp;gt;build_swelling_states()&amp;lt;/code&amp;gt;.&lt;br /&gt;
They remain unchanged in clean as audit evidence and are absent from the&lt;br /&gt;
shared production projection.  This decision excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#250) FEMALE_REPRODUCTIVE_STATES has two incorrect female IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Female IDs &amp;lt;code&amp;gt;OBE&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;UWE&amp;lt;/code&amp;gt; have no production biography match, affecting&lt;br /&gt;
1,404 and 29 rows respectively.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT id, count(*)&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  WHERE id IN (&amp;#039;OBE&amp;#039;, &amp;#039;UWE&amp;#039;)&lt;br /&gt;
  GROUP BY id&lt;br /&gt;
  ORDER BY id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: correct &amp;lt;code&amp;gt;OBE&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;POR&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;UWE&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt; in the clean schema.  The focused clean-stage&lt;br /&gt;
execution changed exactly 1,433 rows; sanity verifies 1,404 &amp;lt;code&amp;gt;POR&amp;lt;/code&amp;gt; and 29 &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;&lt;br /&gt;
rows, no old labels, and applies study-boundary checks afterward.&lt;br /&gt;
&lt;br /&gt;
== * (#251) FEMALE_REPRODUCTIVE_STATES dates fall outside study intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After applying Problem #250&amp;#039;s female corrections, 2,145 dated source rows&lt;br /&gt;
fall before EntryDate or after DepartDate.  This replaces the earlier&lt;br /&gt;
matched-ID-only count of 1,175: correcting &amp;lt;code&amp;gt;OBE&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;POR&amp;lt;/code&amp;gt; adds 677&lt;br /&gt;
before-entry and 293 after-departure rows.  The other corrected ID, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;,&lt;br /&gt;
adds none.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH corrected AS (&lt;br /&gt;
  SELECT source.*&lt;br /&gt;
       , CASE source.id&lt;br /&gt;
           WHEN &amp;#039;OBE&amp;#039; THEN &amp;#039;POR&amp;#039;&lt;br /&gt;
           WHEN &amp;#039;UWE&amp;#039; THEN &amp;#039;MGF&amp;#039;&lt;br /&gt;
           ELSE source.id&lt;br /&gt;
         END AS animid&lt;br /&gt;
    FROM clean.female_reproductive_states AS source)&lt;br /&gt;
SELECT corrected.*&lt;br /&gt;
  FROM corrected&lt;br /&gt;
  JOIN sokwedb.biography_data&lt;br /&gt;
    ON biography_data.animid = corrected.animid&lt;br /&gt;
  WHERE corrected.date &amp;lt; biography_data.entrydate&lt;br /&gt;
        OR corrected.date &amp;gt; biography_data.departdate&lt;br /&gt;
  ORDER BY corrected.animid, corrected.date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: temporarily&lt;br /&gt;
exclude every corrected-female/date row outside the inclusive study interval.&lt;br /&gt;
The shared projection does not alter dates, and sanity independently verifies&lt;br /&gt;
exactly 2,145 exclusions and no overlap with Problems #248 or #254.&lt;br /&gt;
&lt;br /&gt;
== (#252) FEMALE_REPRODUCTIVE_STATES compound offspring IDs represent twins ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Five slash-delimited &amp;lt;code&amp;gt;young_kid_id&amp;lt;/code&amp;gt; values represent twins rather than invalid&lt;br /&gt;
single identifiers.  A single YKID foreign key cannot store both offspring.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT young_kid_id, count(*)&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  WHERE young_kid_id LIKE &amp;#039;%/%&amp;#039;&lt;br /&gt;
  GROUP BY young_kid_id&lt;br /&gt;
  ORDER BY young_kid_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The raw dated counts are &amp;lt;code&amp;gt;GAB2/GAB3&amp;lt;/code&amp;gt; 6, &amp;lt;code&amp;gt;GLI/GLD&amp;lt;/code&amp;gt; 2,013, &amp;lt;code&amp;gt;GY/GL&amp;lt;/code&amp;gt; 296,&lt;br /&gt;
&amp;lt;code&amp;gt;SG/CT&amp;lt;/code&amp;gt; 3,472, and &amp;lt;code&amp;gt;SHO/ROO&amp;lt;/code&amp;gt; 368.  After preceding exclusions, 5,546 source&lt;br /&gt;
rows are eligible and expand to 11,092 target rows; Problem #251 removes 609&lt;br /&gt;
of the &amp;lt;code&amp;gt;SG/CT&amp;lt;/code&amp;gt; rows first.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: split each&lt;br /&gt;
eligible compound row into two target rows, one per offspring, copying all&lt;br /&gt;
other source payload before deriving YKAge separately for each offspring.  The&lt;br /&gt;
projection names only the five approved labels and fails sanity on any new&lt;br /&gt;
slash-delimited label.  It expands 5,546 eligible source rows to 11,092 output&lt;br /&gt;
rows.  See Problem #258 for the resulting female/date/YKID invariant.&lt;br /&gt;
&lt;br /&gt;
== (#253) FEMALE_REPRODUCTIVE_STATES offspring AMA is incorrect ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The unmatched offspring label &amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt; occurs on 287 dated source rows.  The&lt;br /&gt;
investigator confirmed that these values should identify &amp;lt;code&amp;gt;AME&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  WHERE young_kid_id = &amp;#039;AMA&amp;#039;&lt;br /&gt;
  ORDER BY id, date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: correct &amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;AME&amp;lt;/code&amp;gt; in the clean schema.  The focused clean-stage execution changed exactly&lt;br /&gt;
287 rows, and sanity verifies that &amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt; is absent and 287 rows use &amp;lt;code&amp;gt;AME&amp;lt;/code&amp;gt;.&lt;br /&gt;
This correction excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== * (#254) FEMALE_REPRODUCTIVE_STATES offspring TTB2 is unsupported ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The unmatched offspring label &amp;lt;code&amp;gt;TTB2&amp;lt;/code&amp;gt; occurs on 546 dated source rows and has&lt;br /&gt;
no approved production biography identity.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  WHERE young_kid_id = &amp;#039;TTB2&amp;#039;&lt;br /&gt;
  ORDER BY id, date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: temporarily&lt;br /&gt;
exclude exactly these 546 rows.  The projection creates no biography row and&lt;br /&gt;
infers no replacement offspring.  Sanity verifies the exact class and that it&lt;br /&gt;
overlaps neither Problem #248 nor Problem #251.  Problem #248&amp;#039;s 2026-09-24&lt;br /&gt;
amendment converts its 30 rows to a supported code rather than excluding&lt;br /&gt;
them, so this class no longer shares an exclusion mechanism with Problem&lt;br /&gt;
#248, but the two remain non-overlapping in the source data.&lt;br /&gt;
&lt;br /&gt;
== * (#255) Source youngest-offspring ages disagree with biography ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Of 268,336 originally resolved source rows, 115,527 source&lt;br /&gt;
&amp;lt;code&amp;gt;youngest_kid_age&amp;lt;/code&amp;gt; values differ from &amp;lt;code&amp;gt;Date - BIOGRAPHY_DATA.BirthDate&amp;lt;/code&amp;gt;.&lt;br /&gt;
After approved corrections, preceding exclusions, and twin expansion, every&lt;br /&gt;
retained YKID has a biography BirthDate, but 219 nonsplit rows derive a&lt;br /&gt;
negative age from -41 through -1.  Production YKAge must be nonnegative.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH expanded AS (&lt;br /&gt;
  SELECT source.id, source.date&lt;br /&gt;
       , CASE btrim(part.ykid) WHEN &amp;#039;AMA&amp;#039; THEN &amp;#039;AME&amp;#039;&lt;br /&gt;
                              ELSE btrim(part.ykid) END AS ykid&lt;br /&gt;
    FROM clean.female_reproductive_states AS source&lt;br /&gt;
    CROSS JOIN LATERAL&lt;br /&gt;
         regexp_split_to_table(source.young_kid_id, &amp;#039;/&amp;#039;) AS part(ykid)&lt;br /&gt;
   WHERE source.young_kid_id IS NOT NULL&lt;br /&gt;
         AND source.young_kid_id &amp;amp;lt;&amp;amp;gt; &amp;#039;TTB2&amp;#039;)&lt;br /&gt;
SELECT expanded.*, biography_data.birthdate&lt;br /&gt;
     , expanded.date - biography_data.birthdate AS derived_age&lt;br /&gt;
  FROM expanded&lt;br /&gt;
  JOIN sokwedb.biography_data&lt;br /&gt;
    ON biography_data.animid = expanded.ykid&lt;br /&gt;
  WHERE expanded.date - biography_data.birthdate &amp;amp;lt; 0&lt;br /&gt;
  ORDER BY expanded.id, expanded.date, expanded.ykid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The implementation profile must additionally apply Problems #248, #250, and&lt;br /&gt;
#251 in projection order and assert the final 219-row count independently.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: ignore source&lt;br /&gt;
&amp;lt;code&amp;gt;youngest_kid_age&amp;lt;/code&amp;gt; and derive target YKAge from target Date minus the resolved&lt;br /&gt;
offspring&amp;#039;s biography BirthDate.  The projection excludes the 219 rows whose&lt;br /&gt;
derived age is negative.  Sanity verifies zero unresolved retained offspring,&lt;br /&gt;
and validation checks every loaded age against biography.  The original&lt;br /&gt;
115,527-value disagreement remains in the reproducible profile.&lt;br /&gt;
&lt;br /&gt;
== * (#256) FEMALE_REPRO_HISTORY is outside reproductive-state scope ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;clean.female_repro_history&amp;lt;/code&amp;gt; is related historical data but is neither a&lt;br /&gt;
duplicate nor a drop-in supplement for &amp;lt;code&amp;gt;clean.female_reproductive_states&amp;lt;/code&amp;gt;.&lt;br /&gt;
It has 128,306 rows, including 764 female/date keys absent from the requested&lt;br /&gt;
daily source, and overlapping fields differ extensively.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT count(*) AS rows&lt;br /&gt;
     , count(DISTINCT (id, date)) AS female_date_keys&lt;br /&gt;
  FROM clean.female_repro_history;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-21: do not include&lt;br /&gt;
&amp;lt;code&amp;gt;clean.female_repro_history&amp;lt;/code&amp;gt; in this conversion.  Retain the reconciliation&lt;br /&gt;
queries as profile evidence only.  This decision does not exclude any row from&lt;br /&gt;
the requested source.&lt;br /&gt;
&lt;br /&gt;
== (#257) REPRO_STATES AnimID updates bypass integrity checks ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;repro_states_func()&amp;lt;/code&amp;gt; checks sex only on INSERT and checks study boundaries&lt;br /&gt;
only on INSERT or a Date change.  A rollback-only probe changed only AnimID on&lt;br /&gt;
a valid female row to male &amp;lt;code&amp;gt;AL&amp;lt;/code&amp;gt; on 2004-11-19, after AL&amp;#039;s 1999-01-13&lt;br /&gt;
DepartDate; the UPDATE was accepted.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_get_functiondef(&amp;#039;sokwedb.repro_states_func()&amp;#039;::regprocedure);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The returned function gates the sex check on &amp;lt;code&amp;gt;TG_OP = &amp;#039;INSERT&amp;#039;&amp;lt;/code&amp;gt; and the date&lt;br /&gt;
check on &amp;lt;code&amp;gt;TG_OP = &amp;#039;INSERT&amp;#039; OR NEW.date &amp;amp;lt;&amp;amp;gt; OLD.date&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: the owning M4&lt;br /&gt;
reruns female-sex validation whenever AnimID changes and reruns study-boundary&lt;br /&gt;
validation whenever AnimID or Date changes.  Rollback-only tests verified that&lt;br /&gt;
updates to a male AnimID, to a female outside her study interval, and to a date&lt;br /&gt;
outside the current female&amp;#039;s study interval are rejected.&lt;br /&gt;
&lt;br /&gt;
== (#258) REPRO_STATES lacks the female/date/offspring invariant ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production permits duplicate reproductive-state rows.  The earlier proposed&lt;br /&gt;
female/date uniqueness rule is too strict because Problem #252 requires one&lt;br /&gt;
row per twin; the intended identity is female, date, and youngest offspring.&lt;br /&gt;
Rows with no youngest offspring must still be unique per female/date.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT animid, date, ykid, count(*)&lt;br /&gt;
  FROM sokwedb.repro_states&lt;br /&gt;
  GROUP BY animid, date, ykid&lt;br /&gt;
  HAVING count(*) &amp;amp;gt; 1&lt;br /&gt;
  ORDER BY animid, date, ykid NULLS FIRST;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A rollback-only probe inserted two otherwise valid rows for AND on&lt;br /&gt;
2004-11-19 and both were accepted.  The dated approved projection has zero&lt;br /&gt;
duplicate &amp;lt;code&amp;gt;(AnimID, Date, YKID)&amp;lt;/code&amp;gt; keys when NULL YKID is treated as equal.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: a unique index&lt;br /&gt;
was added on &amp;lt;code&amp;gt;(AnimID, Date, YKID) NULLS NOT DISTINCT&amp;lt;/code&amp;gt;, and the table&lt;br /&gt;
documentation describes one row per female/date/youngest-offspring.&lt;br /&gt;
Rollback-only tests verified that duplicate NULL and non-NULL YKID keys are&lt;br /&gt;
rejected while two distinct twin YKIDs on one female/date are accepted.&lt;br /&gt;
&lt;br /&gt;
== (#259) ELO_RANKS_DAILY_KK_F has rows outside GK&amp;#039;s and LB&amp;#039;s recorded community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;clean.elo_ranks_daily___kk_females&amp;lt;/code&amp;gt; maps directly to &amp;lt;code&amp;gt;sokwedb.elo_ranks_daily_kk_f&amp;lt;/code&amp;gt;,&lt;br /&gt;
a table scoped to the &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt; (Kasekela) community.  An inclusive join to&lt;br /&gt;
&amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; found 821 rows for four females that do not fall within a &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt;&lt;br /&gt;
interval on their ranked date:&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;GK&amp;lt;/code&amp;gt;, 181 rows, 1972-11-01 through 1973-04-30: 61 days have no &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt;&lt;br /&gt;
  interval at all, and the remaining 120 days fall within an &amp;lt;code&amp;gt;HK&amp;lt;/code&amp;gt; (Kahama)&lt;br /&gt;
  interval.&lt;br /&gt;
* &amp;lt;code&amp;gt;LB&amp;lt;/code&amp;gt;, 638 rows, 1973-01-01 through 1974-09-30: all 638 days fall within an&lt;br /&gt;
  &amp;lt;code&amp;gt;HK&amp;lt;/code&amp;gt; interval.&lt;br /&gt;
* &amp;lt;code&amp;gt;ML&amp;lt;/code&amp;gt;, 1 row, 1986-10-24: no &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; interval (see Problem #260).&lt;br /&gt;
* &amp;lt;code&amp;gt;AL&amp;lt;/code&amp;gt;, 1 row, 1999-01-14, in &amp;lt;code&amp;gt;sokwedb.elo_ranks_daily_kk_m&amp;lt;/code&amp;gt;: no &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt;&lt;br /&gt;
  interval (see Problem #260).&lt;br /&gt;
&lt;br /&gt;
This spans the historical &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;HK&amp;lt;/code&amp;gt; community split (&amp;lt;code&amp;gt;HK&amp;lt;/code&amp;gt; starts 1973-01-01,&lt;br /&gt;
ends 1977-12-31 per &amp;lt;code&amp;gt;codes.COMM_IDS&amp;lt;/code&amp;gt;).  Neither the target schema nor the&lt;br /&gt;
current loader constrains ranking dates to a &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; interval, so nothing&lt;br /&gt;
in production would reject any treatment of these rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT e.date, e.individual&lt;br /&gt;
  FROM clean.elo_ranks_daily___kk_females e&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
    SELECT 1 FROM sokwedb.comm_membs cm&lt;br /&gt;
     WHERE cm.animid = e.individual&lt;br /&gt;
       AND e.date BETWEEN cm.startdate AND cm.enddate&lt;br /&gt;
       AND cm.commid = &amp;#039;KK&amp;#039;&lt;br /&gt;
  )&lt;br /&gt;
 ORDER BY e.date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Returns 820 rows (&amp;lt;code&amp;gt;GK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;LB&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;ML&amp;lt;/code&amp;gt;); the equivalent query against&lt;br /&gt;
&amp;lt;code&amp;gt;clean.elo_ranks_daily___kk_males&amp;lt;/code&amp;gt; returns 1 row (&amp;lt;code&amp;gt;AL&amp;lt;/code&amp;gt;).&lt;br /&gt;
&amp;lt;code&amp;gt;clean.elo_ranks_daily___mt_females&amp;lt;/code&amp;gt; has zero such rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-22: preserve all 821 rows unchanged.&lt;br /&gt;
The daily Elo ranking is a precomputed, published output; the conversion&amp;#039;s&lt;br /&gt;
role is to carry it forward as supplied, not to re-scope it by community&lt;br /&gt;
after the fact.  No row is excluded and no rank is recalculated.  This&lt;br /&gt;
decision applies identically to the 758 days recorded in &amp;lt;code&amp;gt;HK&amp;lt;/code&amp;gt; and the 63 days&lt;br /&gt;
with no &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; interval at all.&lt;br /&gt;
&lt;br /&gt;
== (#260) ELO_RANKS_DAILY ranking dates trail DepartDate by one day for ML and AL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Two rows are dated exactly one day after the individual&amp;#039;s&lt;br /&gt;
&amp;lt;code&amp;gt;BIOGRAPHY_DATA.DepartDate&amp;lt;/code&amp;gt;:&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;ML&amp;lt;/code&amp;gt; in &amp;lt;code&amp;gt;clean.elo_ranks_daily___kk_females&amp;lt;/code&amp;gt;: ranked 1986-10-24, DepartDate&lt;br /&gt;
  1986-10-23.&lt;br /&gt;
* &amp;lt;code&amp;gt;AL&amp;lt;/code&amp;gt; in &amp;lt;code&amp;gt;clean.elo_ranks_daily___kk_males&amp;lt;/code&amp;gt;: ranked 1999-01-14, DepartDate&lt;br /&gt;
  1999-01-13.&lt;br /&gt;
&lt;br /&gt;
Neither row is otherwise malformed, and the target schema does not constrain&lt;br /&gt;
ranking dates to an individual&amp;#039;s study interval, so both would load without&lt;br /&gt;
error.  Both rows are also members of the 63 &amp;quot;no &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; interval&amp;quot; rows&lt;br /&gt;
described in Problem #259.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT e.date, e.individual, b.departdate&lt;br /&gt;
  FROM clean.elo_ranks_daily___kk_females e&lt;br /&gt;
  JOIN sokwedb.biography_data b ON b.animid = e.individual&lt;br /&gt;
 WHERE e.date = b.departdate + 1;&lt;br /&gt;
-- and the equivalent query against clean.elo_ranks_daily___kk_males&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Returns exactly the &amp;lt;code&amp;gt;ML&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;AL&amp;lt;/code&amp;gt; rows above; no other individual in any of&lt;br /&gt;
the three sources is ranked after their own DepartDate.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-22: preserve both rows unchanged, as&lt;br /&gt;
a final recorded ranking day.  Because each ranked day is a complete &amp;lt;code&amp;gt;1..N&amp;lt;/code&amp;gt;&lt;br /&gt;
permutation, excluding either individual&amp;#039;s single row would leave that day&lt;br /&gt;
looking complete while no longer being one, so no partial-day exclusion is&lt;br /&gt;
created for either row.  This decision was made together with, and is&lt;br /&gt;
consistent with, Problem #259&amp;#039;s community-membership disposition, without&lt;br /&gt;
treating the two as the same underlying question.&lt;br /&gt;
&lt;br /&gt;
== (#261) ELO_RANKS_DAILY documentation misstates the ExpNumBeaten formula ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The shared documentation macro (&amp;lt;code&amp;gt;doc/include/elo_ranks_daily_macro.m4&amp;lt;/code&amp;gt;)&lt;br /&gt;
defines &amp;lt;code&amp;gt;ExpNumBeaten&amp;lt;/code&amp;gt; as &amp;lt;code&amp;gt;EloCardinal&amp;lt;/code&amp;gt; multiplied by the number of&lt;br /&gt;
individuals ranked that day.  Every row in all three sources contradicts that&lt;br /&gt;
formula.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT e.date, e.individual, e.expnumbeaten, e.elocardinal, d.n,&lt;br /&gt;
       e.elocardinal * d.n       AS documented_formula,&lt;br /&gt;
       e.elocardinal * (d.n - 1) AS observed_formula&lt;br /&gt;
  FROM clean.elo_ranks_daily___kk_females e&lt;br /&gt;
  JOIN (SELECT date, count(*) AS n&lt;br /&gt;
          FROM clean.elo_ranks_daily___kk_females&lt;br /&gt;
         GROUP BY date) d ON d.date = e.date&lt;br /&gt;
 ORDER BY abs(e.expnumbeaten - e.elocardinal * d.n) DESC&lt;br /&gt;
 LIMIT 5;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Every row instead matches &amp;lt;code&amp;gt;EloCardinal * (N - 1)&amp;lt;/code&amp;gt;, where &amp;lt;code&amp;gt;N&amp;lt;/code&amp;gt; is the number of&lt;br /&gt;
individuals ranked that day, within 0.00000006 across all three sources --&lt;br /&gt;
i.e. &amp;lt;code&amp;gt;ExpNumBeaten&amp;lt;/code&amp;gt; compares each individual against every other ranked&lt;br /&gt;
individual, not against themselves.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-22: correct&lt;br /&gt;
&amp;lt;code&amp;gt;doc/include/elo_ranks_daily_macro.m4&amp;lt;/code&amp;gt; to describe &amp;lt;code&amp;gt;ExpNumBeaten&amp;lt;/code&amp;gt; as&lt;br /&gt;
&amp;lt;code&amp;gt;EloCardinal * (N - 1)&amp;lt;/code&amp;gt;.  Source &amp;lt;code&amp;gt;ExpNumBeaten&amp;lt;/code&amp;gt; values are copied into&lt;br /&gt;
production unchanged regardless; only the written definition changes, and&lt;br /&gt;
this excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#262) Shared ELO_RANKS_DAILY index macros target only KK_F ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The shared &amp;lt;code&amp;gt;elo_ranks_daily_indexes_create&amp;lt;/code&amp;gt; and&lt;br /&gt;
&amp;lt;code&amp;gt;elo_ranks_daily_indexes_drop&amp;lt;/code&amp;gt; macros ignored their sex and community&lt;br /&gt;
arguments and hard-coded every index name and target relation to&lt;br /&gt;
&amp;lt;code&amp;gt;ELO_RANKS_DAILY_KK_F&amp;lt;/code&amp;gt;.  Consequently, KK_F had its documented unique and&lt;br /&gt;
secondary indexes, while KK_M, MT_F, and MT_M had only their primary-key&lt;br /&gt;
indexes.  In particular, production did not enforce one row per&lt;br /&gt;
&amp;lt;code&amp;gt;(Date, Individual)&amp;lt;/code&amp;gt; on those three tables.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT tablename, indexname, indexdef&lt;br /&gt;
  FROM pg_indexes&lt;br /&gt;
 WHERE schemaname = &amp;#039;sokwedb&amp;#039;&lt;br /&gt;
   AND tablename LIKE &amp;#039;elo_ranks_daily_%&amp;#039;&lt;br /&gt;
 ORDER BY tablename, indexname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Before the repair this returned nine indexes for KK_F and only the primary&lt;br /&gt;
key for each of KK_M, MT_F, and MT_M.  A rollback-only probe confirmed that&lt;br /&gt;
KK_M accepted a duplicate &amp;lt;code&amp;gt;(Date, Individual)&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Implemented on 2026-09-23: both owning macros now derive the relation and&lt;br /&gt;
index names from their sex and community arguments.  All four generated&lt;br /&gt;
create/drop pairs were reviewed independently and installed, leaving each&lt;br /&gt;
table with its primary key plus the expected unique and seven secondary&lt;br /&gt;
indexes.  Rollback-only probes verified that duplicate INSERT and UPDATE&lt;br /&gt;
collisions are rejected in KK_F, KK_M, MT_F, and MT_M, and that all probe&lt;br /&gt;
rows are removed by rollback.&lt;br /&gt;
&lt;br /&gt;
== (#263) BIOGRAPHY_DATA reverse sex validation omits three ELO_RANKS_DAILY tables ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Each ELO_RANKS_DAILY table&amp;#039;s own trigger requires its Individual to have the&lt;br /&gt;
sex encoded by the table suffix.  The reverse &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; trigger,&lt;br /&gt;
however, checked only &amp;lt;code&amp;gt;ELO_RANKS_DAILY_KK_F&amp;lt;/code&amp;gt;.  After a valid row was inserted,&lt;br /&gt;
changing the referenced biography sex could therefore leave KK_M, MT_F, or&lt;br /&gt;
MT_M in a state their own insert/update trigger would reject.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
The defect was reproduced with rollback-only female and male biography&lt;br /&gt;
fixtures.  After inserting one valid rank row in turn, the following updates&lt;br /&gt;
were accepted for MT_F, KK_M, and MT_M, while the equivalent KK_F update was&lt;br /&gt;
rejected:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE sokwedb.biography_data&lt;br /&gt;
   SET sex = &amp;#039;M&amp;#039;&lt;br /&gt;
 WHERE animid = &amp;#039;female fixture&amp;#039;;&lt;br /&gt;
&lt;br /&gt;
UPDATE sokwedb.biography_data&lt;br /&gt;
   SET sex = &amp;#039;F&amp;#039;&lt;br /&gt;
 WHERE animid = &amp;#039;male fixture&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Implemented on 2026-09-23: the owning biography trigger M4 now uses one&lt;br /&gt;
parameterized reverse-check fragment for KK_F, KK_M, MT_F, and MT_M.  Focused&lt;br /&gt;
rollback-only tests verified that female-to-male changes are rejected when&lt;br /&gt;
the individual is referenced by KK_F or MT_F, male-to-female changes are&lt;br /&gt;
rejected when referenced by KK_M or MT_M, and no fixture rows remain.&lt;br /&gt;
&lt;br /&gt;
== * (#264) BIOGRAPHY_DATA REPRO_STATES diagnostic can produce a NULL RAISE option ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
While selecting an existing female for the Problem #263 rollback probe,&lt;br /&gt;
changing &amp;lt;code&amp;gt;AND&amp;lt;/code&amp;gt; from female to male reached the intended reverse REPRO_STATES&lt;br /&gt;
integrity check but failed while constructing its diagnostic.  Several&lt;br /&gt;
REPRO_STATES fields included in &amp;lt;code&amp;gt;DETAIL&amp;lt;/code&amp;gt; are nullable and are concatenated&lt;br /&gt;
without &amp;lt;code&amp;gt;textualize()&amp;lt;/code&amp;gt;.  A NULL field makes the entire &amp;lt;code&amp;gt;DETAIL&amp;lt;/code&amp;gt; expression&lt;br /&gt;
NULL, which PL/pgSQL rejects before it can raise the intended&lt;br /&gt;
&amp;lt;code&amp;gt;integrity_constraint_violation&amp;lt;/code&amp;gt; message.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
BEGIN;&lt;br /&gt;
UPDATE sokwedb.biography_data&lt;br /&gt;
   SET sex = &amp;#039;M&amp;#039;&lt;br /&gt;
 WHERE animid = &amp;#039;AND&amp;#039;;&lt;br /&gt;
ROLLBACK;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This rollback-only probe fails with SQLSTATE &amp;lt;code&amp;gt;22004&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;RAISE statement option&lt;br /&gt;
cannot be null&amp;lt;/code&amp;gt;, in &amp;lt;code&amp;gt;lib.biography_data_func()&amp;lt;/code&amp;gt; instead of reporting the&lt;br /&gt;
referencing REPRO_STATES row.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Unresolved and outside the Elo loader scope.  Use &amp;lt;code&amp;gt;textualize()&amp;lt;/code&amp;gt; or otherwise&lt;br /&gt;
NULL-safe formatting for every nullable REPRO_STATES value included in the&lt;br /&gt;
trigger diagnostic.  The Elo reverse-trigger test uses isolated biography&lt;br /&gt;
fixtures so this pre-existing error path does not mask Problem #263.&lt;br /&gt;
&lt;br /&gt;
== * (#265) BIOGRAPHY_DATA ARRIVALS diagnostic references undeclared a_pid ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
While selecting an existing male for the Problem #263 rollback probe,&lt;br /&gt;
changing &amp;lt;code&amp;gt;AL&amp;lt;/code&amp;gt; from male to female reached the ARRIVALS check for a female&lt;br /&gt;
assigned the male swelling code.  The block&amp;#039;s diagnostic references &amp;lt;code&amp;gt;a_pid&amp;lt;/code&amp;gt;,&lt;br /&gt;
but the block neither declares that variable nor selects &amp;lt;code&amp;gt;roles.pid&amp;lt;/code&amp;gt; into it.&lt;br /&gt;
PostgreSQL therefore fails while parsing the intended diagnostic expression.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
BEGIN;&lt;br /&gt;
UPDATE sokwedb.biography_data&lt;br /&gt;
   SET sex = &amp;#039;F&amp;#039;&lt;br /&gt;
 WHERE animid = &amp;#039;AL&amp;#039;;&lt;br /&gt;
ROLLBACK;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This rollback-only probe fails with SQLSTATE &amp;lt;code&amp;gt;42703&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;column &amp;quot;a_pid&amp;quot; does not&lt;br /&gt;
exist&amp;lt;/code&amp;gt;, in &amp;lt;code&amp;gt;lib.biography_data_func()&amp;lt;/code&amp;gt; instead of reporting the referencing&lt;br /&gt;
ARRIVALS row.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Unresolved and outside the Elo loader scope.  Declare &amp;lt;code&amp;gt;a_pid&amp;lt;/code&amp;gt;, select&lt;br /&gt;
&amp;lt;code&amp;gt;roles.pid&amp;lt;/code&amp;gt; into it in each applicable ARRIVALS query, and retain it in the&lt;br /&gt;
diagnostic.  The Elo reverse-trigger test uses isolated biography fixtures so&lt;br /&gt;
this pre-existing error path does not mask Problem #263.&lt;br /&gt;
&lt;br /&gt;
== * (#266) FERTILITY_FUNC&amp;#039;s overlap check misses contained periods and skips revalidation on Study/AnimID changes ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This is not a data conversion problem.  It is a pre-existing defect in&lt;br /&gt;
production trigger logic, discovered incidentally while working on&lt;br /&gt;
FERTILITY-related conversion problems (see Problems #155 through #158).  The&lt;br /&gt;
conversion-issues log has no better venue for it, so it is recorded here&lt;br /&gt;
rather than left undocumented.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;lib.fertility_func()&amp;lt;/code&amp;gt;, defined in&lt;br /&gt;
&amp;lt;code&amp;gt;db/schemas/lib/triggers/create/fertility.m4&amp;lt;/code&amp;gt;, enforces that, per&lt;br /&gt;
&amp;lt;code&amp;gt;Study&amp;lt;/code&amp;gt; per &amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FERTILITY.StartDate&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;FERTILITY.StopDate&amp;lt;/code&amp;gt;&lt;br /&gt;
periods must not overlap.  The enforcement has two independent gaps.&lt;br /&gt;
&lt;br /&gt;
First, the overlap predicate only tests whether one of &amp;lt;code&amp;gt;NEW&amp;lt;/code&amp;gt;&amp;#039;s own&lt;br /&gt;
endpoints falls inside an existing row&amp;#039;s interval:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
      SELECT fertility.id, fertility.study, fertility.animid&lt;br /&gt;
           , fertility.startdate, fertility.starttype&lt;br /&gt;
           , fertility.stopdate, fertility.stoptype&lt;br /&gt;
        INTO a_id        , a_study        , a_animid&lt;br /&gt;
           , a_startdate        , a_starttype&lt;br /&gt;
           , a_stopdate        , a_stoptype&lt;br /&gt;
        FROM fertility&lt;br /&gt;
        WHERE fertility.id &amp;lt;&amp;gt; NEW.id&lt;br /&gt;
              AND fertility.study = NEW.study&lt;br /&gt;
              AND fertility.animid = NEW.animid&lt;br /&gt;
              AND ((fertility.startdate &amp;lt;= NEW.startdate&lt;br /&gt;
                    AND fertility.stopdate &amp;gt;= NEW.startdate)&lt;br /&gt;
                   OR (fertility.startdate &amp;lt;= NEW.stopdate&lt;br /&gt;
                       AND fertility.stopdate &amp;gt;= NEW.stopdate))&lt;br /&gt;
        -- Produce a consistent error message&lt;br /&gt;
        ORDER BY fertility.id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This misses the case where &amp;lt;code&amp;gt;NEW&amp;lt;/code&amp;gt;&amp;#039;s period fully contains an existing,&lt;br /&gt;
narrower period: neither of &amp;lt;code&amp;gt;NEW.startdate&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;NEW.stopdate&amp;lt;/code&amp;gt; then falls&lt;br /&gt;
inside the existing row, so neither disjunct is true and no overlap is&lt;br /&gt;
reported.  For example, an existing row with &amp;lt;code&amp;gt;StartDate = 2000-03-01&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;StopDate = 2000-03-31&amp;lt;/code&amp;gt; does not conflict, by this test, with a new row&lt;br /&gt;
&amp;lt;code&amp;gt;StartDate = 2000-01-01&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;StopDate = 2000-12-31&amp;lt;/code&amp;gt;, even though the second&lt;br /&gt;
period entirely contains the first.&lt;br /&gt;
&lt;br /&gt;
Second, the trigger only re-runs this check on &amp;lt;code&amp;gt;UPDATE&amp;lt;/code&amp;gt; when&lt;br /&gt;
&amp;lt;code&amp;gt;StartDate&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;StopDate&amp;lt;/code&amp;gt; changes:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  IF TG_OP = &amp;#039;INSERT&amp;#039;&lt;br /&gt;
     OR (NEW.startdate &amp;lt;&amp;gt; OLD.startdate&lt;br /&gt;
         OR NEW.stopdate &amp;lt;&amp;gt; OLD.stopdate) THEN&lt;br /&gt;
    -- Periods of fertility cannot overlap.&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
An &amp;lt;code&amp;gt;UPDATE&amp;lt;/code&amp;gt; that changes &amp;lt;code&amp;gt;Study&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt; instead -- moving a&lt;br /&gt;
period to a different study or a different female -- skips the check&lt;br /&gt;
entirely, even though the row now belongs to a different (&amp;lt;code&amp;gt;Study&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt;) partition that was never checked for overlap against the row&amp;#039;s&lt;br /&gt;
new period.  The comparisons are also NULL-unsafe: &amp;lt;code&amp;gt;&amp;lt;&amp;gt;&amp;lt;/code&amp;gt; evaluates to&lt;br /&gt;
&amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, not &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;, when either side is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, so a change into or&lt;br /&gt;
out of a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;StartDate&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;StopDate&amp;lt;/code&amp;gt; would likewise fail to&lt;br /&gt;
trigger revalidation.&lt;br /&gt;
&lt;br /&gt;
A related, cosmetic defect sits in the same code: the &amp;lt;code&amp;gt;DETAIL&amp;lt;/code&amp;gt; error&lt;br /&gt;
message loses a closing parenthesis after &amp;lt;code&amp;gt;StopType&amp;lt;/code&amp;gt;, so the message&lt;br /&gt;
reads &amp;lt;code&amp;gt;...Value (StopType) = (FOO: Overlapping row has...&amp;lt;/code&amp;gt; instead of&lt;br /&gt;
matching the &amp;lt;code&amp;gt;): Value (...)&amp;lt;/code&amp;gt; pattern used for every other field:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
                       || NEW.stoptype&lt;br /&gt;
                       || &amp;#039;: Overlapping row has Key (FERTILITY.ID) = (&amp;#039;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;doc/src/analyzed/fertility.m4&amp;lt;/code&amp;gt; documents the same incomplete rule the&lt;br /&gt;
code implements, and would need to be corrected alongside any fix:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
For any given female, for any given study, the periods of fertility&lt;br /&gt;
cannot overlap.&lt;br /&gt;
This means that, per female, per study, the |FERTILITY.StopDate| of&lt;br /&gt;
any given row cannot be between, inclusive, the |FERTILITY.StartDate|&lt;br /&gt;
and the |FERTILITY.StopDate| of any other row.&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This prose has the same gap as the code: it defines overlap only in terms of&lt;br /&gt;
one row&amp;#039;s &amp;lt;code&amp;gt;StopDate&amp;lt;/code&amp;gt; falling within another row&amp;#039;s range, so it does not&lt;br /&gt;
describe the containment case either.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Not addressed at this time.  A candidate fix was prepared, verified against&lt;br /&gt;
the standard inclusive-interval-overlap test, but never committed, and is&lt;br /&gt;
recorded here for whoever picks this up next.&lt;br /&gt;
&lt;br /&gt;
In &amp;lt;code&amp;gt;db/schemas/lib/triggers/create/fertility.m4&amp;lt;/code&amp;gt;, replace the overlap&lt;br /&gt;
predicate with the standard two-sided inclusive-interval test, which covers&lt;br /&gt;
containment and shared endpoints in addition to the cases the original&lt;br /&gt;
predicate already caught:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
              AND fertility.startdate &amp;lt;= NEW.stopdate&lt;br /&gt;
              AND NEW.startdate &amp;lt;= fertility.stopdate&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Broaden the &amp;lt;code&amp;gt;UPDATE&amp;lt;/code&amp;gt; revalidation guard to cover every column the&lt;br /&gt;
overlap check depends on, using &amp;lt;code&amp;gt;IS DISTINCT FROM&amp;lt;/code&amp;gt; so a change into or&lt;br /&gt;
out of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; is not silently ignored:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
  IF TG_OP = &amp;#039;INSERT&amp;#039;&lt;br /&gt;
     OR NEW.study IS DISTINCT FROM OLD.study&lt;br /&gt;
     OR NEW.animid IS DISTINCT FROM OLD.animid&lt;br /&gt;
     OR NEW.startdate IS DISTINCT FROM OLD.startdate&lt;br /&gt;
     OR NEW.stopdate IS DISTINCT FROM OLD.stopdate THEN&lt;br /&gt;
    -- Periods of fertility cannot overlap.&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Restore the missing closing parenthesis in the &amp;lt;code&amp;gt;DETAIL&amp;lt;/code&amp;gt; message so it&lt;br /&gt;
reads &amp;lt;code&amp;gt;): Overlapping row has...&amp;lt;/code&amp;gt;, matching every other field in the&lt;br /&gt;
same message.&lt;br /&gt;
&lt;br /&gt;
In &amp;lt;code&amp;gt;doc/src/analyzed/fertility.m4&amp;lt;/code&amp;gt;, correct the prose definition of the&lt;br /&gt;
overlap rule to state the actual (fixed) rule, including the containment and&lt;br /&gt;
shared-endpoint cases:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
For any given female, for any given study, the periods of fertility&lt;br /&gt;
cannot overlap, endpoints included.&lt;br /&gt;
Two rows overlap when the |FERTILITY.StartDate| of each row is on or&lt;br /&gt;
before the |FERTILITY.StopDate| of the other row.&lt;br /&gt;
This includes periods that share an endpoint and periods in which one&lt;br /&gt;
period contains the other.&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Before committing, confirm the corrected predicate does not reject any&lt;br /&gt;
period already present in the loaded production data; if it does, those&lt;br /&gt;
rows are genuine, previously-undetected overlaps in the source and must be&lt;br /&gt;
investigated on their own terms rather than folded into this fix.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=824</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=824"/>
		<updated>2026-09-26T00:20:57Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Adding problems #137 through #265&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commits efa75af037a5e22b2b43675b435f472193a6672f and 921b5ca3e9ca3d1e06fe189b4775aae7dcc484bf&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
9/3/2026 - ICG: Actually this returns cases where a scientific food name is associated with more than one local food name. That is, There can be multiple words for the same latin name.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt; will be derived from&lt;br /&gt;
&amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt;, not &amp;lt;code&amp;gt;fl_sci_food_name&amp;lt;/code&amp;gt;.  Problem&lt;br /&gt;
#87 therefore controls description collisions for this conversion.&lt;br /&gt;
&lt;br /&gt;
The historical query above does not test the relationship stated in the&lt;br /&gt;
heading: it finds a scientific name shared by multiple local names.  Do not use&lt;br /&gt;
the historical count of 47 as a conversion assertion.  If this issue is&lt;br /&gt;
revisited, first replace the query with one that tests the intended&lt;br /&gt;
relationship against the refreshed source.&lt;br /&gt;
&lt;br /&gt;
== (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
9/4/2026 &lt;br /&gt;
Ian fixed mifupa, miti and mizizi in FOOD_BOUT by changing to UNRECORDED. DELETED FROM FOOD_PART_LOOKUP&lt;br /&gt;
Consolidated insects to &amp;#039;dudu&amp;#039;&lt;br /&gt;
changed all &amp;quot;NA&amp;quot; to &amp;quot;None&amp;quot;&lt;br /&gt;
Kept unrecorded&lt;br /&gt;
fixed spellings of utomvi and chipukizi&lt;br /&gt;
&lt;br /&gt;
I made all associated changes in FOOD_BOUT, choosing to use names rather than initials&lt;br /&gt;
&lt;br /&gt;
== (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Use &amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt; for&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt;.  Exclude lookup rows where&lt;br /&gt;
&amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt; is true before deriving descriptions.  Leave each&lt;br /&gt;
noncolliding generalized description unchanged.  When one generalized&lt;br /&gt;
description is shared by multiple eligible local names, derive each target&lt;br /&gt;
description as:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
fl_sci_food_name_gen || &amp;#039; -- &amp;#039; || fl_local_food_name&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This representation is stable and preserves both source values.  Do not use&lt;br /&gt;
mung&amp;#039;s row-order-dependent numeric suffixes.  In the 2026-09-13 snapshot, 39&lt;br /&gt;
generalized descriptions remained shared by 111 eligible local names after&lt;br /&gt;
applying &amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt;.  Rerun a corrected collision query as this&lt;br /&gt;
problem is encountered; the historical count of 42 predates the approved&lt;br /&gt;
exclusion rule.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Do not silently choose between a compound primary part and a conflicting&lt;br /&gt;
explicit second part.  If such a conflict remains after refreshed source&lt;br /&gt;
cleanup, exclude the exact source row and document all source columns in the&lt;br /&gt;
loader predicate.&lt;br /&gt;
&lt;br /&gt;
The dated 2026-09-13 source contained one conflict:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Date !! Focal !! Begin !! End !! Primary part !! Primary name !! Explicit second part !! Second name&lt;br /&gt;
|-&lt;br /&gt;
| 1985-10-26 || EV || 11:22 || 11:51 || MATUNDA; CHIPUKIZI || MBULA || MATUNDA; CHIPUKIZI || BISHURUSHURU&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Under the approved parser, the compound second token is&lt;br /&gt;
&amp;lt;code&amp;gt;CHIPUKIZI&amp;lt;/code&amp;gt;, while the explicit second value is the entire compound&lt;br /&gt;
&amp;lt;code&amp;gt;MATUNDA; CHIPUKIZI&amp;lt;/code&amp;gt;.  The older table above uses the stale spelling&lt;br /&gt;
&amp;lt;code&amp;gt;CHIPUKIZA&amp;lt;/code&amp;gt;.  Reproduce the complete current row exactly before&lt;br /&gt;
adding the exclusion; do not rely on this dated spelling or count.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Add &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt;, described as &amp;lt;code&amp;gt;No data recorded&amp;lt;/code&amp;gt;, to both&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;FOOD_PARTS&amp;lt;/code&amp;gt;.  Create a Seq 2 row when&lt;br /&gt;
either approved secondary component exists.  Use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; only for&lt;br /&gt;
the missing half of that pair:&lt;br /&gt;
&lt;br /&gt;
* part exists, name missing: use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; for FoodName;&lt;br /&gt;
* name exists, part missing: use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; for FoodPart;&lt;br /&gt;
* neither exists: create no Seq 2 row; and&lt;br /&gt;
* both exist: preserve both approved values.&lt;br /&gt;
&lt;br /&gt;
Normalize colon and semicolon delimiters and derive ordered compound tokens for&lt;br /&gt;
Seq 1 and Seq 2.  If a compound-derived second part conflicts with an explicit&lt;br /&gt;
second part, exactly exclude the row under Problem #89.  Do not silently apply&lt;br /&gt;
precedence.  The dated source had 12 secondary names without an explicit second&lt;br /&gt;
part; these are eligible for the approved &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; FoodPart rather&lt;br /&gt;
than exclusion.  Rerun both queries above with the approved clean normalization&lt;br /&gt;
and exclusion projection as this problem is encountered.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, add one deterministic spelling of each case-insensitive GROOM_SCANS extractor value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite gs_extracted_by to the exact spelling stored in clean.people.&lt;br /&gt;
&lt;br /&gt;
The ordinary support-table loader copies these rows to codes.people with active set to true before B-record groom scans are loaded. This preserves all 44,679 source rows and leaves the production foreign key and active-person rule intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of `U` in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-11.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#136) Some GROOM_SCANS direction codes have trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 GROOM_SCANS records where GS_direction is &amp;#039;G &amp;#039; instead of &amp;#039;G&amp;#039;. The trailing Access padding prevents an exact match with the valid one-character direction code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in the clean schema, so query the easy schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;#039;&amp;quot;&amp;#039; || gs_direction || &amp;#039;&amp;quot;&amp;#039; AS untrimmed_direction,&lt;br /&gt;
       &amp;#039;&amp;quot;&amp;#039; || BTRIM(gs_direction) || &amp;#039;&amp;quot;&amp;#039; AS trimmed_direction,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM easy.groom_scans&lt;br /&gt;
 WHERE gs_direction IS DISTINCT FROM BTRIM(gs_direction)&lt;br /&gt;
 GROUP BY gs_direction,&lt;br /&gt;
          BTRIM(gs_direction)&lt;br /&gt;
 ORDER BY gs_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This reports 3 rows having &amp;quot;G &amp;quot;, which normalizes to &amp;quot;G&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Trim leading and trailing whitespace from GS_direction in the clean schema. This preserves all source rows and allows the direction codes to be mapped normally by the production loader.&lt;br /&gt;
&lt;br /&gt;
Resolved by commit 3ad4f5e1e2355895e5e465c4b8049197c8dce295&lt;br /&gt;
&lt;br /&gt;
== * (#137) GROOM_BOUT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A grooming event must relate to a WATCHES row. Normal B-record watches are created from clean.follow, but 1,536 current GROOM_BOUT rows have no follow with the same focal and date after animal-ID whitespace normalization. GROOM_BOUT has no community column, so the loader cannot create a watch directly from the source row.&lt;br /&gt;
&lt;br /&gt;
Of these rows, 1,399 have exactly one COMMUNITY_MEMBERSHIP row covering the observation date. The remaining 137 do not have exactly one dated membership and must not be assigned a community by inference.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid = gb.grm_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = gb.grm_fol_date)&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The temporarily excluded subset is identified by:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid = gb.grm_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = gb.grm_fol_date)&lt;br /&gt;
       AND 1 &amp;lt;&amp;gt; (&lt;br /&gt;
         SELECT count(*)&lt;br /&gt;
           FROM clean.community_membership&lt;br /&gt;
          WHERE community_membership.cm_b_animid =&lt;br /&gt;
                  gb.grm_fol_b_focal_animid&lt;br /&gt;
                AND gb.grm_fol_date BETWEEN&lt;br /&gt;
                      community_membership.cm_start_date&lt;br /&gt;
                      AND community_membership.cm_end_date)&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
As established for ad-hoc observations in Problem #128, prefer an existing B-record WATCHES row and otherwise reuse an existing Other watch for the focal and date. When no watch exists and exactly one COMMUNITY_MEMBERSHIP interval covers the observation date, create an Other watch using that membership&amp;#039;s community. This preserves the grooming bout without falsely representing it as part of a follow.&lt;br /&gt;
&lt;br /&gt;
Do not infer a community from unrelated observations when there is not exactly one dated membership. Temporarily exclude those 137 rows from the production groom-bout loader. Three overlap Problem #138, so this exclusion adds 134 rows to the excluded union. Investigators must determine the focal&amp;#039;s community on the observation date or correct the missing follow or membership data in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#138) GROOM_BOUT initiator or terminator is not a member of the grooming pair ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After applying the approved Problem #112 interpretation of U as unknown and trimming animal IDs under Problem #139, 2,149 GROOM_BOUT rows have a nonempty initiator or terminator animal ID that is neither the focal nor the recorded grooming partner. There are 1,776 initiator mismatches and 2,133 terminator mismatches, with overlap between those sets.&lt;br /&gt;
&lt;br /&gt;
Production GROOMINGS.Initiator and GROOMINGS.Terminator values must reference ROLES.PID rows belonging to participants in the same grooming event. The source does not establish whether the mismatched value, the recorded partner, or another field is incorrect, so the loader cannot safely choose a role or add another participant.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
   gb.grm_fol_b_focal_animid,&lt;br /&gt;
   gb.grm_time_begin,&lt;br /&gt;
   gb.grm_b_partner_animid,&lt;br /&gt;
   gb.grm_direction,&lt;br /&gt;
   invalid_reference.source_column,&lt;br /&gt;
   invalid_reference.animid AS offending_animid,&lt;br /&gt;
   gb.grm_extracted_by,&lt;br /&gt;
   gb.grm_problems,&lt;br /&gt;
   gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
   CROSS JOIN LATERAL (&lt;br /&gt;
     VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
        (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
   ) AS invalid_reference(source_column, animid)&lt;br /&gt;
 WHERE invalid_reference.animid IS NOT NULL&lt;br /&gt;
   AND invalid_reference.animid NOT IN (&lt;br /&gt;
     gb.grm_fol_b_focal_animid,&lt;br /&gt;
     gb.grm_b_partner_animid)&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
      gb.grm_fol_b_focal_animid,&lt;br /&gt;
      gb.grm_time_begin,&lt;br /&gt;
      gb.grm_b_partner_animid,&lt;br /&gt;
      invalid_reference.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows from the production groom-bout loader. The project investigators must determine the correct grooming partner and the correct initiator or terminator in Access. Keep the production requirement that an initiator or terminator reference a participant role in the same grooming event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#139) GROOM_BOUT animal IDs contain edge whitespace ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 50 animal-ID values with leading or trailing whitespace across 49 GROOM_BOUT rows: 36 focal values, 9 partner values, 2 initiator values, and 3 terminator values. The first runtime failure used partner value &amp;lt;code&amp;gt;BE &amp;lt;/code&amp;gt;, which does not reference BIOGRAPHY_DATA even though the trimmed value &amp;lt;code&amp;gt;BE&amp;lt;/code&amp;gt; does.&lt;br /&gt;
&lt;br /&gt;
All 36 focal and all 9 partner values identify existing BIOGRAPHY rows after trimming. Trimming is the established lossless conversion treatment for animal-ID edge whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying easy because the values are normalized in clean&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       animal_id.source_column,&lt;br /&gt;
       animal_id.animid,&lt;br /&gt;
       BTRIM(animal_id.animid) AS trimmed_animid&lt;br /&gt;
  FROM easy.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_FOL_B_focal_AnimId&amp;#039;, gb.grm_fol_b_focal_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_partner_AnimId&amp;#039;, gb.grm_b_partner_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS animal_id(source_column, animid)&lt;br /&gt;
 WHERE animal_id.animid IS DISTINCT FROM BTRIM(animal_id.animid)&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          animal_id.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Trim leading and trailing whitespace from the four GROOM_BOUT animal-ID columns in the clean schema. Apply the Problem #112 conversion of initiator or terminator U values to SQL NULL after trimming. Do not change GRM_direction.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#140) GROOM_BOUT focal or partner animal IDs are absent from BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After the lossless Problem #139 whitespace normalization, 412 GROOM_BOUT rows have a focal or grooming partner that is absent from BIOGRAPHY: 131 rows have an absent focal and 337 have an absent partner. Some rows occur in both counts. Production WATCHES.AnimID and ROLES.Participant values must reference BIOGRAPHY_DATA.AnimID.&lt;br /&gt;
&lt;br /&gt;
The first otherwise eligible runtime failure is the 1992-10-13 PF/MGE bout. The source does not identify a production animal for MGE, so no sentinel or other animal can be substituted safely.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       missing_participant.source_column,&lt;br /&gt;
       missing_participant.animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_FOL_B_focal_AnimId&amp;#039;, gb.grm_fol_b_focal_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_partner_AnimId&amp;#039;, gb.grm_b_partner_animid)&lt;br /&gt;
       ) AS missing_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography&lt;br /&gt;
          WHERE biography.b_animid = missing_participant.animid)&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          missing_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows from the production groom-bout loader. Of the 412 matches, 206 overlap Problems #137 or #138 and 206 are newly excluded. Investigators must identify the intended animals and correct the Access data. Keep the production foreign keys to BIOGRAPHY_DATA intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#141) GROOM_BOUT participants are observed before entry into the study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Seven GROOM_BOUT rows involve a focal or partner before that animal&amp;#039;s BIOGRAPHY.EntryDate: one focal occurrence and six partner occurrences. Production rejects a role dated before its participant entered the study. The first otherwise eligible failure observes SN on 1996-03-11, before its 1996-06-03 entry date.&lt;br /&gt;
&lt;br /&gt;
The identifying predicate uses a strict less-than comparison, so an observation on EntryDate remains valid.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       early_participant.source_column,&lt;br /&gt;
       early_participant.animid,&lt;br /&gt;
       biography.b_entrydate,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_FOL_B_focal_AnimId&amp;#039;, gb.grm_fol_b_focal_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_partner_AnimId&amp;#039;, gb.grm_b_partner_animid)&lt;br /&gt;
       ) AS early_participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography&lt;br /&gt;
         ON biography.b_animid = early_participant.animid&lt;br /&gt;
 WHERE gb.grm_fol_date &amp;lt; biography.b_entrydate&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          early_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows from the production groom-bout loader. One overlaps Problems #137, #138, or #140, so this exclusion adds six rows to the excluded union. Investigators must correct the observation date, participant, or biography entry date in Access. Keep the production rule and its inclusive EntryDate boundary intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#142) GROOM_BOUT participants are observed after departure from the study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Sixty GROOM_BOUT rows involve a focal or partner after that animal&amp;#039;s BIOGRAPHY.DepartDate: five rows have an affected focal and 56 have an affected partner. Some rows occur in both counts. Production rejects a role dated after its participant departed the study. The first otherwise eligible failure observes FI on 1997-07-30, one day after its 1997-07-29 departure date.&lt;br /&gt;
&lt;br /&gt;
The identifying predicate uses a strict greater-than comparison, so an observation on DepartDate remains valid.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       late_participant.source_column,&lt;br /&gt;
       late_participant.animid,&lt;br /&gt;
       biography.b_departdate,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_FOL_B_focal_AnimId&amp;#039;, gb.grm_fol_b_focal_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_partner_AnimId&amp;#039;, gb.grm_b_partner_animid)&lt;br /&gt;
       ) AS late_participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography&lt;br /&gt;
         ON biography.b_animid = late_participant.animid&lt;br /&gt;
 WHERE gb.grm_fol_date &amp;gt; biography.b_departdate&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          late_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows from the production groom-bout loader. Eleven overlap Problems #137, #138, #140, or #141, so this exclusion adds 49 rows to the excluded union. Investigators must correct the observation date, participant, or biography departure date in Access. Keep the production rule and its inclusive DepartDate boundary intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#143) GROOM_BOUT other-partner flags have non-boolean values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
GROOMINGS.Others is a required boolean, while GROOM_BOUT.GRM_other_partners contains several encodings. The established Access boolean-style convention treats NULL or the empty string as false, and the source also contains lowercase y and n values that differ only by case. There are 4,266 blank or NULL values and 561 lowercase values.&lt;br /&gt;
&lt;br /&gt;
After those lossless boundary normalizations, nine rows retain values that cannot be mapped to a boolean: B in three rows, G in one row, M in two rows, and U in three rows. The first otherwise eligible runtime failure has U on the 2002-07-28 SA/SR bout.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary before clean-schema normalization&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(grm_other_partners, &amp;#039;&amp;amp;lt;NULL&amp;amp;gt;&amp;#039;) AS raw_value,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM easy.groom_bout&lt;br /&gt;
 GROUP BY grm_other_partners&lt;br /&gt;
 ORDER BY raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;unmappable values in clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_other_partners,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
 WHERE gb.grm_other_partners NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert NULL and empty GRM_other_partners values to N and normalize lowercase y and n to uppercase. Temporarily exclude the nine rows whose normalized value is not Y or N. One overlaps Problems #137, #138, #140, #141, or #142, so this exclusion adds eight rows to the excluded union.&lt;br /&gt;
&lt;br /&gt;
Investigators must determine whether B, G, M, and U mean that other grooming partners were present and correct the Access data. Keep GROOMINGS.Others required and boolean.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#144) GROOM_BOUT times are outside the production observation window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Six GROOM_BOUT rows have start and stop times outside the production EVENTS range of 04:00 through 20:00 inclusive. All six violate both the start and stop constraints. The first runtime failure is the 2007-03-28 NUR/KS bout from 02:55 through 02:57.&lt;br /&gt;
&lt;br /&gt;
The source does not establish whether these are valid nighttime observations or mistyped times. Changing the production observation window requires a separate decision.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_time_end,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
 WHERE gb.grm_time_begin &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
       OR gb.grm_time_begin &amp;gt; &amp;#039;20:00&amp;#039;::TIME&lt;br /&gt;
       OR (gb.grm_time_end IS NOT NULL&lt;br /&gt;
           AND (gb.grm_time_end &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
                OR gb.grm_time_end &amp;gt; &amp;#039;20:00&amp;#039;::TIME))&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the six rows from the production groom-bout loader. None overlap Problems #137, #138, #140, #141, #142, or #143. Investigators must confirm the intended times or decide separately whether the production observation window should change. Keep the current production constraints intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#145) GROOM_BOUT records the same animal as focal and partner ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Four GROOM_BOUT rows identify the same animal as both the focal and grooming partner. Production represents grooming as a dyadic event and requires each ROLES participant to be unique within an event, so the second role violates the unique participant-and-event constraint. The first failing row is the 2011-07-31 TOM/TOM bout, whose source problems text includes &amp;quot;WRONG FOCAL&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
 WHERE gb.grm_fol_b_focal_animid = gb.grm_b_partner_animid&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the four rows from the production groom-bout loader. None overlap Problems #137, #138, #140, #141, #142, #143, or #144. Investigators must identify the intended focal or grooming partner and correct the Access data. Keep the production requirement that a participant occur at most once in an event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#146) GROOM_SCANS participants are observed after departure from the study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Sixty-four GROOM_SCANS rows involve a grooming participant after that animal&amp;#039;s BIOGRAPHY.DepartDate: eight rows have an affected &amp;lt;code&amp;gt;GS_B_chimp1_AnimId&amp;lt;/code&amp;gt; and 56 have an affected &amp;lt;code&amp;gt;GS_B_chimp2_AnimId&amp;lt;/code&amp;gt;. Production rejects a role dated after its participant departed the study. The first runtime failure observes MT on 1978-05-23, after its 1974-11-01 departure date.&lt;br /&gt;
&lt;br /&gt;
The identifying predicate uses a strict greater-than comparison, so an observation on DepartDate remains valid. Seventeen participant occurrences are on their departure date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       late_participant.source_column,&lt;br /&gt;
       late_participant.animid,&lt;br /&gt;
       biography.b_departdate,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GS_B_chimp1_AnimId&amp;#039;, gs.gs_b_chimp1_animid),&lt;br /&gt;
                (&amp;#039;GS_B_chimp2_AnimId&amp;#039;, gs.gs_b_chimp2_animid)&lt;br /&gt;
       ) AS late_participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography&lt;br /&gt;
         ON biography.b_animid = late_participant.animid&lt;br /&gt;
 WHERE gs.gs_date &amp;gt; biography.b_departdate&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid,&lt;br /&gt;
          late_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the 64 affected rows from the production B-record groom-scan loader. Investigators must correct the observation date, participant, or biography departure date in Access. Keep the production rule and its inclusive DepartDate boundary intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#147) GROOM_SCANS focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A B-record groom scan must relate to a B-type WATCHES row. There are 673 current GROOM_SCANS rows whose focal and date have no matching clean.follow row. Of these, 665 rows have exactly one COMMUNITY_MEMBERSHIP interval covering the observation date and eight rows, comprising three focal/date keys, have no dated membership.&lt;br /&gt;
&lt;br /&gt;
The first runtime failure is the 1978-06-29 ST scan at 14:15. Its source table identifies it as B-record data, and ST has exactly one community membership on that date, so the community is unambiguous even though the corresponding follow is absent.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid = gs.gs_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = gs.gs_date)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The temporarily excluded subset is identified by:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid = gs.gs_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = gs.gs_date)&lt;br /&gt;
       AND 1 &amp;lt;&amp;gt; (&lt;br /&gt;
         SELECT count(*)&lt;br /&gt;
           FROM clean.community_membership&lt;br /&gt;
          WHERE community_membership.cm_b_animid =&lt;br /&gt;
                  gs.gs_fol_b_focal_animid&lt;br /&gt;
                AND gs.gs_date BETWEEN&lt;br /&gt;
                      community_membership.cm_start_date&lt;br /&gt;
                      AND community_membership.cm_end_date)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Reuse an existing B-type WATCHES row when one was created by an earlier loader. Otherwise, when exactly one COMMUNITY_MEMBERSHIP interval covers the observation date, create a B watch using that membership&amp;#039;s community. The GROOM_SCANS source identifies these observations as B-record data, so this does not invent the watch type.&lt;br /&gt;
&lt;br /&gt;
Do not infer a community when no dated membership exists. Temporarily exclude the eight unresolved rows; they do not overlap Problem #146. Investigators must determine the focal&amp;#039;s community on the observation date or correct the missing follow or membership data in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#148) GROOM_SCANS participants are observed before entry into the study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Four GROOM_SCANS rows involve a grooming participant before that animal&amp;#039;s BIOGRAPHY.EntryDate. All four affected values are in &amp;lt;code&amp;gt;GS_B_chimp2_AnimId&amp;lt;/code&amp;gt;. Production rejects a role dated before its participant entered the study. The first runtime failure observes KP on 1981-09-15, before its 1997-03-16 entry date.&lt;br /&gt;
&lt;br /&gt;
The identifying predicate uses a strict less-than comparison, so an observation on EntryDate remains valid. No current GROOM_SCANS participant is observed exactly on its entry date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       early_participant.source_column,&lt;br /&gt;
       early_participant.animid,&lt;br /&gt;
       biography.b_entrydate,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GS_B_chimp1_AnimId&amp;#039;, gs.gs_b_chimp1_animid),&lt;br /&gt;
                (&amp;#039;GS_B_chimp2_AnimId&amp;#039;, gs.gs_b_chimp2_animid)&lt;br /&gt;
       ) AS early_participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography&lt;br /&gt;
         ON biography.b_animid = early_participant.animid&lt;br /&gt;
 WHERE gs.gs_date &amp;lt; biography.b_entrydate&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid,&lt;br /&gt;
          early_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the four affected rows from the production B-record groom-scan loader. They do not overlap Problems #146 or #147. Investigators must correct the observation date, participant, or biography entry date in Access. Keep the production rule and its inclusive EntryDate boundary intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#149) GROOM_SCANS times are outside the production observation window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Sixty-one GROOM_SCANS rows have times outside the production EVENTS range of 04:00 through 20:00 inclusive. The values range from 00:00 through 23:50. The first runtime failure is the 1982-01-13 PI scan at 03:30.&lt;br /&gt;
&lt;br /&gt;
The source does not establish whether these are valid nighttime observations or mistyped times. Changing the production observation window requires a separate decision. No current GROOM_SCANS row occurs exactly at 04:00 or 20:00.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
   gs.gs_fol_b_focal_animid,&lt;br /&gt;
   gs.gs_time,&lt;br /&gt;
   gs.gs_b_chimp1_animid,&lt;br /&gt;
   gs.gs_b_chimp2_animid,&lt;br /&gt;
   gs.gs_direction,&lt;br /&gt;
   gs.gs_extracted_by,&lt;br /&gt;
   gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE gs.gs_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
   OR gs.gs_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
      gs.gs_fol_b_focal_animid,&lt;br /&gt;
      gs.gs_time,&lt;br /&gt;
      gs.gs_b_chimp1_animid,&lt;br /&gt;
      gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the 61 affected rows from the production B-record groom-scan loader. They do not overlap Problems #146, #147, or #148. Investigators must confirm the intended times or decide separately whether the production observation window should change. Keep the current production constraints intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#150) GROOM_SCANS unknown direction cannot use the production UNKPair role ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Thirty-one GROOM_SCANS rows have &amp;lt;code&amp;gt;GS_direction = &amp;#039;U&amp;#039;&amp;lt;/code&amp;gt;. The loader maps this source code to paired &amp;lt;code&amp;gt;UNKPair&amp;lt;/code&amp;gt; roles, which represent a dyadic interaction whose direction is unknown. The current production &amp;lt;code&amp;gt;roles_func&amp;lt;/code&amp;gt; and EVENTS documentation permit Actor, Actee, and Mutual roles for GSCAN events but reject &amp;lt;code&amp;gt;UNKPair&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The first runtime failure is the 1982-02-07 ST scan at 11:55 with participants FF and GB. Its source comment says &amp;quot;NO DIRECTION INDICATED ON TIKI OR ON B RECORDS&amp;quot;. None of the 31 rows overlaps Problems #146 through #149.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE gs.gs_direction = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the 31 source rows while retaining the current production role restriction. The investigator does not want structural schema changes in the B-record loader commit. Flag permitting paired &amp;lt;code&amp;gt;UNKPair&amp;lt;/code&amp;gt; roles for GSCAN events, with the corresponding EVENTS documentation change, as a structural problem for the next iteration.&lt;br /&gt;
&lt;br /&gt;
Do not convert &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; to Mutual because unknown direction does not establish symmetric grooming.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#151) GROOM_SCANS participants are absent from BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Two hundred twenty-five GROOM_SCANS rows have a grooming participant absent from BIOGRAPHY: 33 affected values are in &amp;lt;code&amp;gt;GS_B_chimp1_AnimId&amp;lt;/code&amp;gt; and 192 are in &amp;lt;code&amp;gt;GS_B_chimp2_AnimId&amp;lt;/code&amp;gt;. Production ROLES.Participant must reference BIOGRAPHY_DATA.AnimID. The first runtime failure is the 1982-03-09 FF scan at 15:45, whose second participant is TP.&lt;br /&gt;
&lt;br /&gt;
The source does not identify a production animal that can safely replace an absent participant.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       missing_participant.source_column,&lt;br /&gt;
       missing_participant.animid,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GS_B_chimp1_AnimId&amp;#039;, gs.gs_b_chimp1_animid),&lt;br /&gt;
                (&amp;#039;GS_B_chimp2_AnimId&amp;#039;, gs.gs_b_chimp2_animid)&lt;br /&gt;
       ) AS missing_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography&lt;br /&gt;
          WHERE biography.b_animid = missing_participant.animid)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid,&lt;br /&gt;
          missing_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the 225 affected rows from the production B-record groom-scan loader. They do not overlap Problems #146 through #150. Investigators must identify the intended animals and correct the Access data. Keep the production foreign key to BIOGRAPHY_DATA intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#152) GROOM_SCANS direction values cannot be mapped to production roles ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After the Problem #136 whitespace normalization, 48 GROOM_SCANS rows have a direction other than G, R, M, or U: 43 are NULL, one is F, one is H, and three are T. The first runtime failure that reached the direction mapping has H on the 1989-06-22 GB scan at 08:15.&lt;br /&gt;
&lt;br /&gt;
The source does not establish which participant groomed the other, whether grooming was mutual, or whether direction was unknown. The loader cannot choose production roles safely.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE COALESCE(gs.gs_direction, &amp;#039;&amp;#039;) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the 48 affected rows from the production B-record groom-scan loader. Three overlap Problems #146 through #151, so this exclusion adds 45 rows to the excluded union. Investigators must determine the intended direction and correct the Access data. Do not infer roles from participant order or map missing direction to Mutual or unknown without a documented decision.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#153) GROOM_SCANS records the same animal as both participants ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Twelve GROOM_SCANS rows identify the same animal in &amp;lt;code&amp;gt;GS_B_chimp1_AnimId&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;GS_B_chimp2_AnimId&amp;lt;/code&amp;gt;. Production represents groom scans as dyadic events and requires each ROLES participant to be unique within an event. The first runtime failure is the 2015-03-16 NUR/NUR scan at 13:00, whose source comment says &amp;quot;Same ID for chimp1 and chimp2&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE gs.gs_b_chimp1_animid = gs.gs_b_chimp2_animid&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the 12 affected rows from the production B-record groom-scan loader. They do not overlap Problems #146 through #152. Investigators must identify the intended first or second participant, or decide separately how self-grooming should be represented. Keep the production requirement that a participant occur at most once in an event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#154) GROOM_SCANS records have no observation time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Eight GROOM_SCANS rows have NULL &amp;lt;code&amp;gt;GS_time&amp;lt;/code&amp;gt; values. Production EVENTS requires a time, and all eight source comments say &amp;quot;no time&amp;quot;. None of the rows overlaps Problems #146 through #153.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       gs.gs_direction,&lt;br /&gt;
       gs.gs_extracted_by,&lt;br /&gt;
       gs.gs_comments&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE gs.gs_time IS NULL&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the eight affected rows from the production B-record groom-scan loader. Investigators must recover the observation times from another source or document a separate policy for missing event times. Do not invent a time from the scan date, record order, or neighboring observations.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#155) FERTILITY boundary-type codes need production support rows ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The 168 &amp;lt;code&amp;gt;clean.fertility&amp;lt;/code&amp;gt; rows use four &amp;lt;code&amp;gt;StartType&amp;lt;/code&amp;gt; values and four &amp;lt;code&amp;gt;StopType&amp;lt;/code&amp;gt;&lt;br /&gt;
values. Production FERTILITY requires these values to reference&lt;br /&gt;
FERTILITY_STARTS and FERTILITY_STOPS, but both support tables are empty.&lt;br /&gt;
&lt;br /&gt;
The fertility domains exactly match the BIOGRAPHY_DATA boundary-type domains:&lt;br /&gt;
&amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;C&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;I&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;O&amp;lt;/code&amp;gt; match ENTRYTYPES, while &amp;lt;code&amp;gt;D&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;E&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;O&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;P&amp;lt;/code&amp;gt; match&lt;br /&gt;
DEPARTTYPES.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;#039;StartType&amp;#039; AS source_column,&lt;br /&gt;
       fertility.starttype AS code,&lt;br /&gt;
       COUNT(*) AS rows&lt;br /&gt;
  FROM clean.fertility AS fertility&lt;br /&gt;
  GROUP BY fertility.starttype&lt;br /&gt;
UNION ALL&lt;br /&gt;
SELECT &amp;#039;StopType&amp;#039;,&lt;br /&gt;
       fertility.stoptype,&lt;br /&gt;
       COUNT(*)&lt;br /&gt;
  FROM clean.fertility AS fertility&lt;br /&gt;
  GROUP BY fertility.stoptype&lt;br /&gt;
ORDER BY source_column, code;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed 2026-09-13 source contains StartType &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt; (70 rows), &amp;lt;code&amp;gt;C&amp;lt;/code&amp;gt; (56),&lt;br /&gt;
&amp;lt;code&amp;gt;I&amp;lt;/code&amp;gt; (34), and &amp;lt;code&amp;gt;O&amp;lt;/code&amp;gt; (8), and StopType &amp;lt;code&amp;gt;D&amp;lt;/code&amp;gt; (74), &amp;lt;code&amp;gt;E&amp;lt;/code&amp;gt; (1), &amp;lt;code&amp;gt;O&amp;lt;/code&amp;gt; (75), and &amp;lt;code&amp;gt;P&amp;lt;/code&amp;gt;&lt;br /&gt;
(18).&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The project approved using the corresponding ENTRYTYPES and DEPARTTYPES&lt;br /&gt;
meanings for fertility boundary types. Populate FERTILITY_STARTS with &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;&lt;br /&gt;
(Birth), &amp;lt;code&amp;gt;C&amp;lt;/code&amp;gt; (Start of confirmed AnimID), &amp;lt;code&amp;gt;I&amp;lt;/code&amp;gt; (Immigration), and &amp;lt;code&amp;gt;O&amp;lt;/code&amp;gt;&lt;br /&gt;
(Initiation of close observation). Populate FERTILITY_STOPS with &amp;lt;code&amp;gt;D&amp;lt;/code&amp;gt; (Death),&lt;br /&gt;
&amp;lt;code&amp;gt;E&amp;lt;/code&amp;gt; (Emigration), &amp;lt;code&amp;gt;O&amp;lt;/code&amp;gt; (End of observation; Present in the most recent census),&lt;br /&gt;
and &amp;lt;code&amp;gt;P&amp;lt;/code&amp;gt; (Permanent disappearance).&lt;br /&gt;
&lt;br /&gt;
This is a support-table decision and excludes no source rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#156) FERTILITY source intervals require ordered dates ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production FERTILITY does not have a table constraint requiring &amp;lt;code&amp;gt;StartDate &amp;lt;=&lt;br /&gt;
StopDate&amp;lt;/code&amp;gt;. Loading a reversed source interval would therefore preserve an&lt;br /&gt;
invalid period without necessarily producing a database error.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fertility.studyid,&lt;br /&gt;
       fertility.animalid,&lt;br /&gt;
       fertility.startdate,&lt;br /&gt;
       fertility.starttype,&lt;br /&gt;
       fertility.stopdate,&lt;br /&gt;
       fertility.stoptype&lt;br /&gt;
  FROM clean.fertility AS fertility&lt;br /&gt;
  WHERE fertility.startdate &amp;gt; fertility.stopdate&lt;br /&gt;
  ORDER BY fertility.studyid,&lt;br /&gt;
           fertility.animalid,&lt;br /&gt;
           fertility.startdate,&lt;br /&gt;
           fertility.stopdate;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed 2026-09-13 source contains zero reversed intervals.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Enforce &amp;lt;code&amp;gt;StartDate &amp;lt;= StopDate&amp;lt;/code&amp;gt; in the read-only fertility conversion sanity&lt;br /&gt;
gate. Do not add a production table constraint as part of this conversion.&lt;br /&gt;
Abort before writing any production rows if a future source snapshot contains&lt;br /&gt;
a reversed interval.&lt;br /&gt;
&lt;br /&gt;
This decision excludes no source rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#157) Selected FERTILITY periods cross BIOGRAPHY_DATA boundaries ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Nine &amp;lt;code&amp;gt;clean.fertility&amp;lt;/code&amp;gt; rows extend beyond one or more BIOGRAPHY_DATA date&lt;br /&gt;
boundaries. Production does not require fertility periods to remain within&lt;br /&gt;
BirthDate, EntryDate, and DepartDate, so the database does not establish&lt;br /&gt;
whether these differences are valid history or source errors.&lt;br /&gt;
&lt;br /&gt;
Investigators reviewed the nine rows and selected six for temporary exclusion:&lt;br /&gt;
AR, BH, GLIB1, MAM, NV, and RAF. The remaining boundary differences for SAF,&lt;br /&gt;
SHO, and SN1 are approved for conversion and are not excluded.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader; all six source fields identify each approved exclusion&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fertility.studyid,&lt;br /&gt;
       fertility.animalid,&lt;br /&gt;
       fertility.startdate,&lt;br /&gt;
       fertility.starttype,&lt;br /&gt;
       fertility.stopdate,&lt;br /&gt;
       fertility.stoptype&lt;br /&gt;
  FROM clean.fertility AS fertility&lt;br /&gt;
  WHERE (fertility.studyid,&lt;br /&gt;
         fertility.animalid,&lt;br /&gt;
         fertility.startdate,&lt;br /&gt;
         fertility.starttype,&lt;br /&gt;
         fertility.stopdate,&lt;br /&gt;
         fertility.stoptype)&lt;br /&gt;
        IN ((4, &amp;#039;AR&amp;#039;, DATE &amp;#039;1985-07-11&amp;#039;, &amp;#039;B&amp;#039;, DATE &amp;#039;1987-07-19&amp;#039;, &amp;#039;D&amp;#039;),&lt;br /&gt;
            (4, &amp;#039;BH&amp;#039;, DATE &amp;#039;1971-07-05&amp;#039;, &amp;#039;B&amp;#039;, DATE &amp;#039;1971-09-20&amp;#039;, &amp;#039;D&amp;#039;),&lt;br /&gt;
            (4, &amp;#039;GLIB1&amp;#039;, DATE &amp;#039;2011-07-11&amp;#039;, &amp;#039;B&amp;#039;, DATE &amp;#039;2012-01-21&amp;#039;, &amp;#039;D&amp;#039;),&lt;br /&gt;
            (4, &amp;#039;MAM&amp;#039;, DATE &amp;#039;2004-02-12&amp;#039;, &amp;#039;B&amp;#039;, DATE &amp;#039;2011-12-07&amp;#039;, &amp;#039;D&amp;#039;),&lt;br /&gt;
            (4, &amp;#039;NV&amp;#039;, DATE &amp;#039;1965-09-15&amp;#039;, &amp;#039;I&amp;#039;, DATE &amp;#039;1975-03-12&amp;#039;, &amp;#039;D&amp;#039;),&lt;br /&gt;
            (4, &amp;#039;RAF&amp;#039;, DATE &amp;#039;1986-10-21&amp;#039;, &amp;#039;C&amp;#039;, DATE &amp;#039;1996-03-27&amp;#039;, &amp;#039;D&amp;#039;))&lt;br /&gt;
  ORDER BY fertility.animalid,&lt;br /&gt;
           fertility.startdate;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The query returns exactly six rows. Because this is a new FERTILITY loader,&lt;br /&gt;
none overlaps an earlier source-row exclusion. The distinct exclusion union is&lt;br /&gt;
six rows, leaving 162 eligible rows from the refreshed 168-row source.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the six identified rows from the production FERTILITY&lt;br /&gt;
loader. Keep the full six-column predicate adjacent to its diagnostic query so&lt;br /&gt;
the exclusion cannot broaden silently. Investigators must correct the Access&lt;br /&gt;
fertility or biography dates, or approve a later interpretation that permits&lt;br /&gt;
these periods, before the rows are restored.&lt;br /&gt;
&lt;br /&gt;
Do not exclude SAF, SHO, or SN1. Their boundary differences were reviewed and&lt;br /&gt;
approved for conversion.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#158) FERTILITY StudyId 4 has no production STUDIES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All 168 &amp;lt;code&amp;gt;clean.fertility&amp;lt;/code&amp;gt; rows have integer &amp;lt;code&amp;gt;StudyId&amp;lt;/code&amp;gt; 4. Production&lt;br /&gt;
FERTILITY.Study is text and must reference STUDIES, but the follow-derived&lt;br /&gt;
support data contains no Study code &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt;. The Access snapshot contains no study&lt;br /&gt;
lookup table, description, foreign key, or imported metadata that gives StudyId&lt;br /&gt;
4 a textual name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fertility.studyid,&lt;br /&gt;
       COUNT(*) AS rows&lt;br /&gt;
  FROM clean.fertility AS fertility&lt;br /&gt;
  GROUP BY fertility.studyid&lt;br /&gt;
  ORDER BY fertility.studyid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed 2026-09-13 source returns only StudyId 4, with 168 rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Preserve the source value without inventing unsupported semantics. Add Study&lt;br /&gt;
code &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt; to STUDIES with description &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt;, and map &amp;lt;code&amp;gt;clean.fertility.studyid&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;FERTILITY.Study&amp;lt;/code&amp;gt; using &amp;lt;code&amp;gt;studyid::text&amp;lt;/code&amp;gt; at the load boundary.&lt;br /&gt;
&lt;br /&gt;
This support and representation decision excludes no source rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#159) FOOD_LOOKUP status flags need explicit conversion semantics ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;clean.food_lookup&amp;lt;/code&amp;gt; contains &amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt; and&lt;br /&gt;
&amp;lt;code&amp;gt;FL_unverified&amp;lt;/code&amp;gt; flags, but the mung support and event loaders ignored&lt;br /&gt;
both.  Loading all lookup rows would admit names explicitly marked for&lt;br /&gt;
exclusion; treating unverified names the same way would discard data without&lt;br /&gt;
investigator approval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;The following is a dated diagnostic. Rerun it against the refreshed clean&lt;br /&gt;
schema and record exact referring rows before implementing exclusions.&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH normalized_lookup AS (&lt;br /&gt;
  SELECT UPPER(BTRIM(fl_local_food_name)) AS foodname,&lt;br /&gt;
         fl_exclude,&lt;br /&gt;
         &amp;quot;FL_unverified&amp;quot; AS fl_unverified&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
),&lt;br /&gt;
used AS (&lt;br /&gt;
  SELECT &amp;#039;seq1&amp;#039; AS slot,&lt;br /&gt;
         fb_fol_date,&lt;br /&gt;
         fb_fol_b_focal_animid,&lt;br /&gt;
         fb_begin_feed_time,&lt;br /&gt;
         fb_end_feed_time,&lt;br /&gt;
         UPPER(BTRIM(fb_fl_local_food_name)) AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;seq2&amp;#039;,&lt;br /&gt;
         fb_fol_date,&lt;br /&gt;
         fb_fol_b_focal_animid,&lt;br /&gt;
         fb_begin_feed_time,&lt;br /&gt;
         fb_end_feed_time,&lt;br /&gt;
         UPPER(BTRIM(fb_local_food_name2))&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    WHERE NULLIF(BTRIM(fb_local_food_name2), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT used.*,&lt;br /&gt;
       normalized_lookup.fl_exclude,&lt;br /&gt;
       normalized_lookup.fl_unverified&lt;br /&gt;
  FROM used&lt;br /&gt;
  JOIN normalized_lookup USING (foodname)&lt;br /&gt;
  WHERE normalized_lookup.fl_exclude IS TRUE&lt;br /&gt;
     OR normalized_lookup.fl_unverified IS TRUE&lt;br /&gt;
  ORDER BY used.fb_fol_date,&lt;br /&gt;
           used.fb_fol_b_focal_animid,&lt;br /&gt;
           used.fb_begin_feed_time,&lt;br /&gt;
           used.slot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-13 snapshot had 18 excluded lookup rows and 359 unverified lookup&lt;br /&gt;
rows.  Excluded names were referenced by 538 primary and 24 secondary bout&lt;br /&gt;
items, 562 occurrences total.  Unverified names were referenced by 728 bout&lt;br /&gt;
items.  These sets may overlap and must be profiled separately.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;FL_exclude = true&amp;lt;/code&amp;gt; disqualifies the lookup entry.  Do not load that&lt;br /&gt;
name into &amp;lt;code&amp;gt;FOOD_NAMES&amp;lt;/code&amp;gt;, and exactly exclude every referring food-bout&lt;br /&gt;
item after approved case and edge-whitespace normalization.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;FL_unverified = true&amp;lt;/code&amp;gt; does not disqualify a lookup entry or&lt;br /&gt;
referring bout.  Preserve the flag in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt;, document that it has no&lt;br /&gt;
production destination, and load the otherwise eligible name and bout.&lt;br /&gt;
&lt;br /&gt;
The sanity gate must prove that no excluded name enters production and that no&lt;br /&gt;
unverified name is omitted merely because it is unverified.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#160) Converted food events require a consumer ROLES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The mung food loader inserted &amp;lt;code&amp;gt;EVENTS&amp;lt;/code&amp;gt; and&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_EVENTS&amp;lt;/code&amp;gt; rows but inserted no &amp;lt;code&amp;gt;ROLES&amp;lt;/code&amp;gt; row.  Production&lt;br /&gt;
documentation says the food event&amp;#039;s role identifies the individual consuming&lt;br /&gt;
the food.  The roles trigger permits at most one role and requires that&lt;br /&gt;
participant to equal the related watch focal.  Although absence currently&lt;br /&gt;
produces a warning rather than a hard constraint failure, omitting the consumer&lt;br /&gt;
would make the conversion incomplete.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Run after a rollback-only or disposable-database load.&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT events.eid,&lt;br /&gt;
       watches.wid,&lt;br /&gt;
       watches.animid AS focal,&lt;br /&gt;
       COUNT(roles.pid) AS role_rows,&lt;br /&gt;
       MIN(roles.role) AS role,&lt;br /&gt;
       MIN(roles.participant) AS participant&lt;br /&gt;
  FROM sokwedb.events&lt;br /&gt;
  JOIN sokwedb.watches&lt;br /&gt;
    ON watches.wid = events.wid&lt;br /&gt;
  LEFT JOIN sokwedb.roles&lt;br /&gt;
    ON roles.eid = events.eid&lt;br /&gt;
  WHERE events.behavior = &amp;#039;FOOD&amp;#039;&lt;br /&gt;
  GROUP BY events.eid,&lt;br /&gt;
           watches.wid,&lt;br /&gt;
           watches.animid&lt;br /&gt;
  HAVING COUNT(roles.pid) &amp;lt;&amp;gt; 1&lt;br /&gt;
      OR MIN(roles.role) &amp;lt;&amp;gt; &amp;#039;Solo&amp;#039;&lt;br /&gt;
      OR MIN(roles.participant) &amp;lt;&amp;gt; watches.animid&lt;br /&gt;
  ORDER BY events.eid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For every eligible food bout, insert exactly one &amp;lt;code&amp;gt;ROLES&amp;lt;/code&amp;gt; row related&lt;br /&gt;
to the newly created event.  Use role code &amp;lt;code&amp;gt;Solo&amp;lt;/code&amp;gt; and the selected B&lt;br /&gt;
watch&amp;#039;s exact &amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt; as Participant.  Capture the generated EID;&lt;br /&gt;
do not infer event identity from sequence state or source ordering.&lt;br /&gt;
&lt;br /&gt;
This decision excludes no source rows.  The loader and parity checks must prove&lt;br /&gt;
one food event, one Solo consumer role, and one or two contiguous food-detail&lt;br /&gt;
rows per eligible source bout.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#161) FOOD_VARIATIONS conversion is deferred ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;clean.food_variations_lookup&amp;lt;/code&amp;gt; has a corresponding production&lt;br /&gt;
&amp;lt;code&amp;gt;housekeeping.FOOD_VARIATIONS&amp;lt;/code&amp;gt; table, but the mung branch never&lt;br /&gt;
loaded it.  The current food-event conversion requires decisions for missing&lt;br /&gt;
and unresolved LocalName values that are independent of loading food bouts and&lt;br /&gt;
their two support tables.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;This is a dated scope diagnostic, not an exclusion query.&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) AS source_rows,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE NULLIF(BTRIM(fvl_food_spelling_variant), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
       ) AS missing_variants,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE NULLIF(BTRIM(fvl_fl_local_food_name), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
       ) AS missing_local_names&lt;br /&gt;
  FROM clean.food_variations_lookup;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-13 snapshot contained 1,575 rows, no blank variants, 224 blank or&lt;br /&gt;
NULL local names, and 38 additional nonblank local names without a normalized&lt;br /&gt;
match in &amp;lt;code&amp;gt;clean.food_lookup&amp;lt;/code&amp;gt;.  Rerun and extend the query in the&lt;br /&gt;
future variations conversion; do not reuse these counts as exclusions.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Explicitly defer &amp;lt;code&amp;gt;clean.food_variations_lookup&amp;lt;/code&amp;gt; to a separate future&lt;br /&gt;
conversion.  The present food conversion must not insert&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_VARIATIONS&amp;lt;/code&amp;gt; rows, must not depend on that table to normalize&lt;br /&gt;
bout names, and must not claim variation parity.&lt;br /&gt;
&lt;br /&gt;
Preserve this issue and the source table for the future conversion.  The&lt;br /&gt;
deferral excludes no food-bout rows and does not authorize dropping variation&lt;br /&gt;
data.&lt;br /&gt;
&lt;br /&gt;
== (#162) Eligible FOOD_LOOKUP rows lack generalized scientific names ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The approved &amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt; source is&lt;br /&gt;
&amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt;, but 360 eligible lookup rows have a blank or&lt;br /&gt;
NULL value.  All 360 names are used by food bouts, in 812 primary and 111&lt;br /&gt;
secondary items.  Of these names, 359 are marked &amp;lt;code&amp;gt;FL_unverified&amp;lt;/code&amp;gt; and&lt;br /&gt;
must not be disqualified under Problem #159.  The target description is NOT&lt;br /&gt;
NULL, nonempty, and case-insensitively unique.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-13: use&lt;br /&gt;
&amp;lt;code&amp;gt;Unknown -- &amp;amp;lt;local name&amp;amp;gt;&amp;lt;/code&amp;gt; for an eligible blank generalized name.&lt;br /&gt;
This stable value identifies the exact source name and does not consult&lt;br /&gt;
&amp;lt;code&amp;gt;fl_sci_food_name&amp;lt;/code&amp;gt;.  Nonblank descriptions continue to follow&lt;br /&gt;
Problem #87.  This mapping excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#163) FOOD_BOUT duration disagrees with elapsed event time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The refreshed source has 1,136 rows where &amp;lt;code&amp;gt;fb_duration&amp;lt;/code&amp;gt; differs from&lt;br /&gt;
the exact minutes between begin and end.  Production stores the original begin&lt;br /&gt;
and end times but has no independent duration column.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not alter event times to reproduce &amp;lt;code&amp;gt;fb_duration&amp;lt;/code&amp;gt;.  Retain the&lt;br /&gt;
field in &amp;lt;code&amp;gt;clean.food_bout&amp;lt;/code&amp;gt;, emit a sanity warning with the refreshed&lt;br /&gt;
count, and document that it has no production destination.  This issue causes&lt;br /&gt;
no exclusions.&lt;br /&gt;
&lt;br /&gt;
== (#164) FOOD_BOUT update metadata has no production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The refreshed source has 17,702 rows with non-NULL &amp;lt;code&amp;gt;fb_update&amp;lt;/code&amp;gt;.&lt;br /&gt;
None of the food target tables has a corresponding update-metadata column.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Retain &amp;lt;code&amp;gt;fb_update&amp;lt;/code&amp;gt; in &amp;lt;code&amp;gt;clean.food_bout&amp;lt;/code&amp;gt;, emit a sanity&lt;br /&gt;
warning with the refreshed count, and do not overload event notes or another&lt;br /&gt;
production field.  This issue causes no exclusions.&lt;br /&gt;
&lt;br /&gt;
== * (#165) ATTENDANCE contains exact duplicate rows ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The source table has no primary key and contains values that are identical in&lt;br /&gt;
all 17 columns.  Loading every copy would create indistinguishable production&lt;br /&gt;
events and repeated sequence values.  Retaining an arbitrary copy would not be&lt;br /&gt;
an exact source-row policy because no source value distinguishes the copies.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Run against the refreshed &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  The grouping explicitly&lt;br /&gt;
compares all 17 typed source columns.&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*,&lt;br /&gt;
       COUNT(*) AS copies&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  GROUP BY attendance.a_date,&lt;br /&gt;
           attendance.a_cl_community_id,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num,&lt;br /&gt;
           attendance.a_type_of_cycle,&lt;br /&gt;
           attendance.a_observer1_id,&lt;br /&gt;
           attendance.a_observer2_id,&lt;br /&gt;
           attendance.a_bananas_given,&lt;br /&gt;
           attendance.a_degree_of_arrival,&lt;br /&gt;
           attendance.a_degree_of_departure,&lt;br /&gt;
           attendance.a_time_start,&lt;br /&gt;
           attendance.a_time_end,&lt;br /&gt;
           attendance.a_duration_of_obs,&lt;br /&gt;
           attendance.day,&lt;br /&gt;
           attendance.mo,&lt;br /&gt;
           attendance.yr,&lt;br /&gt;
           attendance.cycle_old&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num,&lt;br /&gt;
           attendance.a_time_start;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot contained 207 duplicate values, each occurring twice:&lt;br /&gt;
414 source rows total.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exclude every copy of each exact&lt;br /&gt;
duplicate value.  Do not use &amp;lt;code&amp;gt;DISTINCT&amp;lt;/code&amp;gt; to retain one arbitrary copy.&lt;br /&gt;
The shared attendance projection and loader must include this query beside the&lt;br /&gt;
Problem #165 predicate and prove that every row whose complete typed value has&lt;br /&gt;
multiplicity greater than one is excluded.&lt;br /&gt;
&lt;br /&gt;
Implemented in &amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt; as&lt;br /&gt;
&amp;lt;code&amp;gt;duplicate_count &amp;amp;gt; 1&amp;lt;/code&amp;gt;, partitioned over all 17 columns.  On&lt;br /&gt;
2026-09-15 the investigator approved semantic validation: the dated 207 values&lt;br /&gt;
and 414 rows are audit evidence, not frozen acceptance criteria.  Sanity&lt;br /&gt;
asserts that every refreshed duplicate copy is classified and excluded.&lt;br /&gt;
Remove only when exact duplicates leave the source.&lt;br /&gt;
&lt;br /&gt;
== * (#166) ATTENDANCE sequence values are not contiguous per animal and date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production attendance sequence values are expected to be exactly&lt;br /&gt;
&amp;lt;code&amp;gt;1..N&amp;lt;/code&amp;gt; for each animal/date.  Exact duplicates and other targeted&lt;br /&gt;
row exclusions can themselves create duplicate values or gaps, so sequence&lt;br /&gt;
validity must be checked after all other attendance exclusions.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;This query expresses the approved staging against the 2026-09-14 schema.&lt;br /&gt;
Keep it synchronized with Problems #165 and #168 through #178 when predicates&lt;br /&gt;
change.&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source AS (&lt;br /&gt;
  SELECT row_number() OVER () AS source_id,&lt;br /&gt;
         attendance.*,&lt;br /&gt;
         COUNT(*) OVER (&lt;br /&gt;
           PARTITION BY attendance.a_date,&lt;br /&gt;
                        attendance.a_cl_community_id,&lt;br /&gt;
                        attendance.a_b_animid,&lt;br /&gt;
                        attendance.a_seq_num,&lt;br /&gt;
                        attendance.a_type_of_cycle,&lt;br /&gt;
                        attendance.a_observer1_id,&lt;br /&gt;
                        attendance.a_observer2_id,&lt;br /&gt;
                        attendance.a_bananas_given,&lt;br /&gt;
                        attendance.a_degree_of_arrival,&lt;br /&gt;
                        attendance.a_degree_of_departure,&lt;br /&gt;
                        attendance.a_time_start,&lt;br /&gt;
                        attendance.a_time_end,&lt;br /&gt;
                        attendance.a_duration_of_obs,&lt;br /&gt;
                        attendance.day,&lt;br /&gt;
                        attendance.mo,&lt;br /&gt;
                        attendance.yr,&lt;br /&gt;
                        attendance.cycle_old&lt;br /&gt;
         ) AS duplicate_count&lt;br /&gt;
    FROM clean.attendance AS attendance&lt;br /&gt;
),&lt;br /&gt;
independent_eligible AS (&lt;br /&gt;
  SELECT source.*&lt;br /&gt;
    FROM source&lt;br /&gt;
      JOIN sokwedb.biography_data AS biography&lt;br /&gt;
        ON biography.animid = source.a_b_animid::text&lt;br /&gt;
    WHERE source.duplicate_count = 1&lt;br /&gt;
      AND source.a_date BETWEEN biography.entrydate AND biography.departdate&lt;br /&gt;
      AND COALESCE(BTRIM(source.a_observer1_id::text), &amp;#039;NONE&amp;#039;)&lt;br /&gt;
            &amp;amp;lt;&amp;amp;gt; COALESCE(BTRIM(source.a_observer2_id::text), &amp;#039;NONE&amp;#039;)&lt;br /&gt;
      AND BTRIM(source.a_type_of_cycle::text)&lt;br /&gt;
            NOT IN (&amp;#039;-&amp;#039;, &amp;#039;0.30&amp;#039;, &amp;#039;0.33&amp;#039;)&lt;br /&gt;
      AND source.cycle_old IS NOT NULL&lt;br /&gt;
      AND source.cycle_old::text = BTRIM(source.cycle_old::text)&lt;br /&gt;
      AND (source.a_bananas_given IS NULL&lt;br /&gt;
           OR source.a_bananas_given = TRUNC(source.a_bananas_given))&lt;br /&gt;
      AND source.a_time_start IS NOT NULL&lt;br /&gt;
      AND source.a_time_end IS NOT NULL&lt;br /&gt;
      AND source.a_time_start &amp;amp;lt;= source.a_time_end&lt;br /&gt;
      AND source.a_duration_of_obs IS NOT DISTINCT FROM&lt;br /&gt;
            EXTRACT(EPOCH FROM&lt;br /&gt;
                    (source.a_time_end - source.a_time_start)) / 60&lt;br /&gt;
      AND source.day IS NOT DISTINCT FROM&lt;br /&gt;
        EXTRACT(DAY FROM source.a_date)::integer&lt;br /&gt;
      AND source.mo IS NOT DISTINCT FROM&lt;br /&gt;
        EXTRACT(MONTH FROM source.a_date)::integer&lt;br /&gt;
      AND source.yr IS NOT DISTINCT FROM&lt;br /&gt;
        EXTRACT(YEAR FROM source.a_date)::integer&lt;br /&gt;
),&lt;br /&gt;
overlap_rows AS (&lt;br /&gt;
  SELECT DISTINCT first_row.source_id&lt;br /&gt;
    FROM independent_eligible AS first_row&lt;br /&gt;
      JOIN independent_eligible AS second_row&lt;br /&gt;
        ON second_row.source_id &amp;amp;lt;&amp;amp;gt; first_row.source_id&lt;br /&gt;
       AND second_row.a_date = first_row.a_date&lt;br /&gt;
       AND second_row.a_b_animid = first_row.a_b_animid&lt;br /&gt;
       AND second_row.a_time_start &amp;amp;lt;= first_row.a_time_end&lt;br /&gt;
       AND second_row.a_time_end &amp;amp;gt;= first_row.a_time_start&lt;br /&gt;
),&lt;br /&gt;
eligible_before_seq AS (&lt;br /&gt;
  SELECT independent_eligible.*&lt;br /&gt;
    FROM independent_eligible&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
      SELECT 1&lt;br /&gt;
        FROM overlap_rows&lt;br /&gt;
          WHERE overlap_rows.source_id = independent_eligible.source_id)&lt;br /&gt;
),&lt;br /&gt;
bad_groups AS (&lt;br /&gt;
  SELECT a_date,&lt;br /&gt;
         a_b_animid&lt;br /&gt;
    FROM eligible_before_seq&lt;br /&gt;
    GROUP BY a_date,&lt;br /&gt;
             a_b_animid&lt;br /&gt;
    HAVING MIN(a_seq_num) IS NULL&lt;br /&gt;
       OR MIN(a_seq_num) &amp;amp;lt;&amp;amp;gt; 1&lt;br /&gt;
       OR MAX(a_seq_num) IS DISTINCT FROM COUNT(*)&lt;br /&gt;
       OR COUNT(DISTINCT a_seq_num) &amp;amp;lt;&amp;amp;gt; COUNT(*)&lt;br /&gt;
)&lt;br /&gt;
SELECT eligible_before_seq.*&lt;br /&gt;
  FROM eligible_before_seq&lt;br /&gt;
    JOIN bad_groups&lt;br /&gt;
      USING (a_date, a_b_animid)&lt;br /&gt;
  ORDER BY eligible_before_seq.a_date,&lt;br /&gt;
           eligible_before_seq.a_b_animid,&lt;br /&gt;
           eligible_before_seq.a_seq_num,&lt;br /&gt;
           eligible_before_seq.a_time_start;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed projection contained 453 invalid animal/date groups and 910 rows&lt;br /&gt;
after the other approved exclusions and the Problem #171 mapping.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exclude every row in each invalid&lt;br /&gt;
animal/date group.  Do not renumber, retain a partial group, or select rows by&lt;br /&gt;
order.  Compute Problem #166 last, after Problems #165, #168 through #170, and&lt;br /&gt;
#172 through #178,&lt;br /&gt;
and include the complete identification query beside the loader predicate.&lt;br /&gt;
Sanity must prove every remaining group has non-NULL sequence values exactly&lt;br /&gt;
equal to &amp;lt;code&amp;gt;1..N&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Implemented last in &amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt; by&lt;br /&gt;
&amp;lt;code&amp;gt;attendance_bad_seq_groups&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;problem_166&amp;lt;/code&amp;gt;.  The&lt;br /&gt;
dated 453 groups and 910 rows are audit evidence.  Sanity proves every&lt;br /&gt;
refreshed invalid group is completely excluded and every eligible group is&lt;br /&gt;
exactly &amp;lt;code&amp;gt;1..N&amp;lt;/code&amp;gt;.  Remove only when all post-exclusion groups are&lt;br /&gt;
contiguous.&lt;br /&gt;
&lt;br /&gt;
== (#167) ATTENDANCE observer values require PEOPLE support rows ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The observer columns contain hundreds of case and edge-whitespace variants.&lt;br /&gt;
Observer codes must be matched by their standardized value so variants such as&lt;br /&gt;
&amp;lt;code&amp;gt;Mike&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;mike&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;Mike &amp;lt;/code&amp;gt; resolve to one&lt;br /&gt;
&amp;lt;code&amp;gt;PEOPLE.Person&amp;lt;/code&amp;gt;.  Standardized values absent from&lt;br /&gt;
&amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; require support rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH raw_observers AS (&lt;br /&gt;
  SELECT a_observer1_id::text AS raw_person&lt;br /&gt;
    FROM clean.attendance&lt;br /&gt;
    WHERE a_observer1_id IS NOT NULL&lt;br /&gt;
  UNION&lt;br /&gt;
  SELECT a_observer2_id::text&lt;br /&gt;
    FROM clean.attendance&lt;br /&gt;
    WHERE a_observer2_id IS NOT NULL&lt;br /&gt;
), standardized AS (&lt;br /&gt;
  SELECT raw_person,&lt;br /&gt;
         LOWER(NORMALIZE(BTRIM(raw_person))) AS standard_person&lt;br /&gt;
    FROM raw_observers&lt;br /&gt;
)&lt;br /&gt;
SELECT standardized.standard_person,&lt;br /&gt;
       ARRAY_AGG(standardized.raw_person&lt;br /&gt;
                 ORDER BY standardized.raw_person) AS raw_spellings,&lt;br /&gt;
       ARRAY_AGG(people.person ORDER BY people.person)&lt;br /&gt;
         FILTER (WHERE people.person IS NOT NULL) AS existing_codes&lt;br /&gt;
  FROM standardized&lt;br /&gt;
  LEFT JOIN codes.people&lt;br /&gt;
    ON LOWER(NORMALIZE(BTRIM(people.person))) =&lt;br /&gt;
         standardized.standard_person&lt;br /&gt;
  GROUP BY standardized.standard_person&lt;br /&gt;
  ORDER BY standardized.standard_person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed 2026-09-14 snapshot had 932 distinct raw observer spellings in&lt;br /&gt;
718 standardized classes.  Of those classes, 245 matched an existing&lt;br /&gt;
&amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; row and 473 required a new support row.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: standardize observer identifiers&lt;br /&gt;
with &amp;lt;code&amp;gt;LOWER(NORMALIZE(BTRIM(value)))&amp;lt;/code&amp;gt;.  Resolve a class to its existing&lt;br /&gt;
&amp;lt;code&amp;gt;PEOPLE.Person&amp;lt;/code&amp;gt; spelling when present.  For an unmatched class, choose&lt;br /&gt;
one deterministic trimmed source spelling, preferring mixed case and then&lt;br /&gt;
bytewise order, and add it as an active support row.  Use that spelling for&lt;br /&gt;
Person and Name and &amp;lt;code&amp;gt;Attendance observer code: &amp;amp;lt;code&amp;amp;gt;&amp;lt;/code&amp;gt; for&lt;br /&gt;
Description.  Map NULL to the existing &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; code.  This support&lt;br /&gt;
decision excludes no rows except where Problem #168 also applies.&lt;br /&gt;
&lt;br /&gt;
Implemented by &amp;lt;code&amp;gt;attendance_observer_codes&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt; and the ordered&lt;br /&gt;
&amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; insert in &amp;lt;code&amp;gt;conversion/load_attendance.sql&amp;lt;/code&amp;gt;.  Sanity&lt;br /&gt;
pins all 718 canonical codes using an explicitly byte-ordered, length-prefixed&lt;br /&gt;
serialization, all 932 raw-to-canonical mappings, and 473 missing support rows.&lt;br /&gt;
Remove only when every standardized observer value already resolves to one&lt;br /&gt;
active row.&lt;br /&gt;
&lt;br /&gt;
== * (#168) ATTENDANCE Recorder and Observer2 resolve to the same person code ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;ARRIVALS_A&amp;lt;/code&amp;gt; requires Recorder and Observer2 to differ.  Some source&lt;br /&gt;
rows contain observer spellings that become equal after standardizing case and&lt;br /&gt;
edge whitespace; rows with both observers NULL also become equal after the&lt;br /&gt;
approved NULL-to-&amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; mapping.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE LOWER(NORMALIZE(COALESCE(&lt;br /&gt;
          BTRIM(attendance.a_observer1_id::text), &amp;#039;NONE&amp;#039;))) =&lt;br /&gt;
        LOWER(NORMALIZE(COALESCE(&lt;br /&gt;
          BTRIM(attendance.a_observer2_id::text), &amp;#039;NONE&amp;#039;)))&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num,&lt;br /&gt;
           attendance.a_time_start;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned 1,544 rows: 1,484 with equal trimmed&lt;br /&gt;
non-NULL spellings, 39 additional case-equivalent pairs, and 21 with both&lt;br /&gt;
observers NULL.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Do not fabricate or discard an observer.  Include the query beside the&lt;br /&gt;
Problem #168 predicate and prove no eligible detail has equal observer codes.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_168&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 1,544 rows are&lt;br /&gt;
audit evidence; sanity proves that no refreshed eligible row has equal derived&lt;br /&gt;
observer codes.  Remove only when no derived Recorder equals Observer2.&lt;br /&gt;
&lt;br /&gt;
== * (#169) ATTENDANCE focal does not resolve to BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An attendance watch and its Solo role require an exact production AnimID.&lt;br /&gt;
Thousands of source rows use values absent from &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt;,&lt;br /&gt;
mostly &amp;lt;code&amp;gt;NYANI&amp;lt;/code&amp;gt;.  Case or whitespace repair is not approved.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
      FROM sokwedb.biography_data AS biography&lt;br /&gt;
      WHERE biography.animid = attendance.a_b_animid::text)&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num,&lt;br /&gt;
           attendance.a_time_start;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The exact-match query returned 6,877 rows on 2026-09-14.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Do not trim, case-fold, or map an unmatched focal to a sentinel.  Include&lt;br /&gt;
the query beside the Problem #169 predicate and prove every eligible focal&lt;br /&gt;
resolves exactly.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_169&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 6,877 rows are&lt;br /&gt;
audit evidence; sanity proves that every refreshed eligible focal resolves.&lt;br /&gt;
Remove only when every exact source focal resolves.&lt;br /&gt;
&lt;br /&gt;
== * (#170) ATTENDANCE focal is outside its study dates ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A role participant must be under study on the event date.  Some exact-matching&lt;br /&gt;
focals have attendance dates before EntryDate or after DepartDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*,&lt;br /&gt;
       biography.entrydate,&lt;br /&gt;
       biography.departdate&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
    JOIN sokwedb.biography_data AS biography&lt;br /&gt;
      ON biography.animid = attendance.a_b_animid::text&lt;br /&gt;
  WHERE attendance.a_date NOT BETWEEN&lt;br /&gt;
        biography.entrydate AND biography.departdate&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned six rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Include the query beside the Problem #170 predicate and prove every&lt;br /&gt;
eligible participant is under study on the event date.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_170&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated six rows are audit&lt;br /&gt;
evidence; sanity proves that every refreshed eligible focal is under study on&lt;br /&gt;
its attendance date.  Remove only when every exact focal is under study.&lt;br /&gt;
&lt;br /&gt;
== (#171) ATTENDANCE direction fields contain the -9 sentinel ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production degrees may be NULL or integers from 0 through 359.  The source uses&lt;br /&gt;
&amp;lt;code&amp;gt;-9&amp;lt;/code&amp;gt; in one or both direction fields to mean that the direction was&lt;br /&gt;
not seen.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE COALESCE(attendance.a_degree_of_arrival = -9, FALSE)&lt;br /&gt;
     OR COALESCE(attendance.a_degree_of_departure = -9, FALSE)&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned 7,228 rows.  Individual arrival and departure&lt;br /&gt;
counts overlap and must not be summed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Confirmed by the project investigators on 2026-09-15: translate each&lt;br /&gt;
&amp;lt;code&amp;gt;-9&amp;lt;/code&amp;gt; independently to NULL because it means the direction was not&lt;br /&gt;
seen.  Preserve existing NULL and values from 0 through 359 exactly.  Keep this&lt;br /&gt;
query beside the Problem #171 predicate to profile the mapped source condition,&lt;br /&gt;
but do not add Problem #171 to the exclusion array.  Prove every eligible&lt;br /&gt;
target degree equals &amp;lt;code&amp;gt;NULLIF(source_degree, -9)&amp;lt;/code&amp;gt; and satisfies the&lt;br /&gt;
target domain.&lt;br /&gt;
&lt;br /&gt;
Implemented as the profiled &amp;lt;code&amp;gt;problem_171&amp;lt;/code&amp;gt; condition and the&lt;br /&gt;
&amp;lt;code&amp;gt;NULLIF(..., -9)&amp;lt;/code&amp;gt; arrival/departure projection in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 7,228 rows are&lt;br /&gt;
audit evidence; no rows are excluded solely by this issue.  Sanity proves each&lt;br /&gt;
refreshed eligible target equals &amp;lt;code&amp;gt;NULLIF(source_degree, -9)&amp;lt;/code&amp;gt;.  Remove&lt;br /&gt;
the mapping only when neither source degree contains &amp;lt;code&amp;gt;-9&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== * (#172) ATTENDANCE swelling values have no approved CYCLE_STATES mapping ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Most source swelling spellings have an exact lossless representation in&lt;br /&gt;
&amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt;.  The values &amp;lt;code&amp;gt;-&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.30&amp;lt;/code&amp;gt;, and&lt;br /&gt;
&amp;lt;code&amp;gt;0.33&amp;lt;/code&amp;gt; do not.  Mapping &amp;lt;code&amp;gt;-&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; or rounding&lt;br /&gt;
the numeric values would change source meaning.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE BTRIM(attendance.a_type_of_cycle::text)&lt;br /&gt;
        IN (&amp;#039;-&amp;#039;, &amp;#039;0.30&amp;#039;, &amp;#039;0.33&amp;#039;)&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned 3,238 rows: 3,221 with &amp;lt;code&amp;gt;-&amp;lt;/code&amp;gt;, two with&lt;br /&gt;
&amp;lt;code&amp;gt;0.30&amp;lt;/code&amp;gt;, and 15 with &amp;lt;code&amp;gt;0.33&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Do not map or round these values.  The explicit supported projection is:&lt;br /&gt;
&amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;0.00&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;;&lt;br /&gt;
&amp;lt;code&amp;gt;.25&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;;&lt;br /&gt;
&amp;lt;code&amp;gt;.5&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;0.50&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;;&lt;br /&gt;
&amp;lt;code&amp;gt;.75&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;;&lt;br /&gt;
&amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;1.00&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt;;&lt;br /&gt;
&amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;N/A&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;; and&lt;br /&gt;
&amp;lt;code&amp;gt;u&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;.  Include the query beside the&lt;br /&gt;
Problem #172 predicate and reject any refreshed spelling outside this closed&lt;br /&gt;
mapping.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_172&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 3,238 rows are&lt;br /&gt;
audit evidence; sanity rejects any refreshed eligible spelling outside the&lt;br /&gt;
closed exact mapping.  Remove only when every source spelling belongs to it.&lt;br /&gt;
&lt;br /&gt;
== * (#173) ATTENDANCE CycleOld is missing or has edge spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;ARRIVALS_A.CycleOld&amp;lt;/code&amp;gt; is required and may not contain only spaces.&lt;br /&gt;
The source has NULL values and edge-spaced values.  Trimming would alter the&lt;br /&gt;
initial digitization recorded by this field.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE attendance.cycle_old IS NULL&lt;br /&gt;
     OR attendance.cycle_old::text &amp;amp;lt;&amp;amp;gt;&lt;br /&gt;
        BTRIM(attendance.cycle_old::text)&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned 417 rows: 229 NULL and 188 edge-spaced values.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Preserve all eligible CycleOld text exactly; do not trim or invent a&lt;br /&gt;
missing value.  Include the query beside the Problem #173 predicate.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_173&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 417 rows are audit&lt;br /&gt;
evidence; sanity proves every refreshed eligible CycleOld is non-NULL and&lt;br /&gt;
edge-space free.  Remove only when every source value satisfies that rule.&lt;br /&gt;
&lt;br /&gt;
== * (#174) ATTENDANCE bananas contains a fractional value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;ARRIVALS_A.Bananas&amp;lt;/code&amp;gt; is an integer.  One source value is fractional,&lt;br /&gt;
and rounding or truncating it would alter the recorded value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE attendance.a_bananas_given IS NOT NULL&lt;br /&gt;
    AND attendance.a_bananas_given &amp;amp;lt;&amp;amp;gt;&lt;br /&gt;
        TRUNC(attendance.a_bananas_given)&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned one row with value &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude the returned row.&lt;br /&gt;
Do not round or truncate it.  Include the query beside the Problem #174&lt;br /&gt;
predicate and prove all eligible banana values are integral and in range.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_174&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated one-row result is&lt;br /&gt;
audit evidence; sanity proves all refreshed eligible banana values are&lt;br /&gt;
integral and in range.  Remove only when all source values are integral.&lt;br /&gt;
&lt;br /&gt;
== * (#175) ATTENDANCE event interval is missing or inverted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Attendance event Start and Stop are required, and Start may not be after Stop.&lt;br /&gt;
One source row has no Start and other rows have Start after Stop.  Production&lt;br /&gt;
attendance events cannot represent an inferred cross-midnight interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE attendance.a_time_start IS NULL&lt;br /&gt;
      OR attendance.a_time_end IS NULL&lt;br /&gt;
     OR attendance.a_time_start &amp;amp;gt; attendance.a_time_end&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned 28 rows: one missing Start, no missing Stop,&lt;br /&gt;
and 27 inverted intervals.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Do not swap times, infer dates, derive Start from duration, or clamp a&lt;br /&gt;
value.  Include the query beside the Problem #175 predicate.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_175&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 28 rows are audit&lt;br /&gt;
evidence; sanity proves every refreshed eligible interval has both endpoints&lt;br /&gt;
and is ordered.  Remove only when all source intervals satisfy that rule.&lt;br /&gt;
&lt;br /&gt;
== * (#176) ATTENDANCE duration disagrees with the event interval ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
For otherwise ordered intervals, some stored durations differ from the exact&lt;br /&gt;
minutes between Start and Stop.  Production has no independent attendance&lt;br /&gt;
duration destination, so loading only the times would discard a conflicting&lt;br /&gt;
source assertion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*,&lt;br /&gt;
       EXTRACT(EPOCH FROM&lt;br /&gt;
               (attendance.a_time_end - attendance.a_time_start)) / 60&lt;br /&gt;
         AS computed_minutes&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
  WHERE attendance.a_time_start IS NOT NULL&lt;br /&gt;
    AND attendance.a_time_start &amp;amp;lt;= attendance.a_time_end&lt;br /&gt;
    AND attendance.a_duration_of_obs IS DISTINCT FROM&lt;br /&gt;
        EXTRACT(EPOCH FROM&lt;br /&gt;
                (attendance.a_time_end - attendance.a_time_start)) / 60&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned 15 rows, separate from Problem #175.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Do not alter event times or discard the disagreement as warning-only.&lt;br /&gt;
Include the query beside the Problem #176 predicate and prove exact agreement&lt;br /&gt;
for every eligible row.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_176&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 15 rows are audit&lt;br /&gt;
evidence; sanity proves every refreshed eligible duration agrees exactly with&lt;br /&gt;
elapsed time.  Remove only when every otherwise valid source row agrees.&lt;br /&gt;
&lt;br /&gt;
== * (#177) ATTENDANCE intervals overlap for one animal and date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production warns when one animal is recorded at the feeding station in&lt;br /&gt;
overlapping intervals.  Exact duplicates and independently invalid rows must&lt;br /&gt;
be removed first so this issue identifies only otherwise eligible overlap&lt;br /&gt;
participants.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;This query is intentionally complete.  Keep its independent eligibility&lt;br /&gt;
predicates synchronized with Problems #165 and #168 through #176/#178.&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source AS (&lt;br /&gt;
  SELECT row_number() OVER () AS source_id,&lt;br /&gt;
         attendance.*,&lt;br /&gt;
         COUNT(*) OVER (&lt;br /&gt;
           PARTITION BY attendance.a_date,&lt;br /&gt;
                        attendance.a_cl_community_id,&lt;br /&gt;
                        attendance.a_b_animid,&lt;br /&gt;
                        attendance.a_seq_num,&lt;br /&gt;
                        attendance.a_type_of_cycle,&lt;br /&gt;
                        attendance.a_observer1_id,&lt;br /&gt;
                        attendance.a_observer2_id,&lt;br /&gt;
                        attendance.a_bananas_given,&lt;br /&gt;
                        attendance.a_degree_of_arrival,&lt;br /&gt;
                        attendance.a_degree_of_departure,&lt;br /&gt;
                        attendance.a_time_start,&lt;br /&gt;
                        attendance.a_time_end,&lt;br /&gt;
                        attendance.a_duration_of_obs,&lt;br /&gt;
                        attendance.day,&lt;br /&gt;
                        attendance.mo,&lt;br /&gt;
                        attendance.yr,&lt;br /&gt;
                        attendance.cycle_old&lt;br /&gt;
         ) AS duplicate_count&lt;br /&gt;
    FROM clean.attendance AS attendance&lt;br /&gt;
),&lt;br /&gt;
independent_eligible AS (&lt;br /&gt;
  SELECT source.*&lt;br /&gt;
    FROM source&lt;br /&gt;
      JOIN sokwedb.biography_data AS biography&lt;br /&gt;
        ON biography.animid = source.a_b_animid::text&lt;br /&gt;
    WHERE source.duplicate_count = 1&lt;br /&gt;
      AND source.a_date BETWEEN biography.entrydate AND biography.departdate&lt;br /&gt;
      AND COALESCE(BTRIM(source.a_observer1_id::text), &amp;#039;NONE&amp;#039;)&lt;br /&gt;
            &amp;amp;lt;&amp;amp;gt; COALESCE(BTRIM(source.a_observer2_id::text), &amp;#039;NONE&amp;#039;)&lt;br /&gt;
      AND BTRIM(source.a_type_of_cycle::text)&lt;br /&gt;
            NOT IN (&amp;#039;-&amp;#039;, &amp;#039;0.30&amp;#039;, &amp;#039;0.33&amp;#039;)&lt;br /&gt;
      AND source.cycle_old IS NOT NULL&lt;br /&gt;
      AND source.cycle_old::text = BTRIM(source.cycle_old::text)&lt;br /&gt;
      AND (source.a_bananas_given IS NULL&lt;br /&gt;
           OR source.a_bananas_given = TRUNC(source.a_bananas_given))&lt;br /&gt;
      AND source.a_time_start IS NOT NULL&lt;br /&gt;
      AND source.a_time_end IS NOT NULL&lt;br /&gt;
      AND source.a_time_start &amp;amp;lt;= source.a_time_end&lt;br /&gt;
      AND source.a_duration_of_obs IS NOT DISTINCT FROM&lt;br /&gt;
            EXTRACT(EPOCH FROM&lt;br /&gt;
                    (source.a_time_end - source.a_time_start)) / 60&lt;br /&gt;
      AND source.day IS NOT DISTINCT FROM&lt;br /&gt;
        EXTRACT(DAY FROM source.a_date)::integer&lt;br /&gt;
      AND source.mo IS NOT DISTINCT FROM&lt;br /&gt;
        EXTRACT(MONTH FROM source.a_date)::integer&lt;br /&gt;
      AND source.yr IS NOT DISTINCT FROM&lt;br /&gt;
        EXTRACT(YEAR FROM source.a_date)::integer&lt;br /&gt;
)&lt;br /&gt;
SELECT DISTINCT first_row.*&lt;br /&gt;
  FROM independent_eligible AS first_row&lt;br /&gt;
    JOIN independent_eligible AS second_row&lt;br /&gt;
      ON second_row.source_id &amp;amp;lt;&amp;amp;gt; first_row.source_id&lt;br /&gt;
     AND second_row.a_date = first_row.a_date&lt;br /&gt;
     AND second_row.a_b_animid = first_row.a_b_animid&lt;br /&gt;
     AND second_row.a_time_start &amp;amp;lt;= first_row.a_time_end&lt;br /&gt;
     AND second_row.a_time_end &amp;amp;gt;= first_row.a_time_start&lt;br /&gt;
  ORDER BY first_row.a_date,&lt;br /&gt;
           first_row.a_b_animid,&lt;br /&gt;
           first_row.a_time_start,&lt;br /&gt;
           first_row.a_time_end;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed projection returned 428 otherwise eligible source rows after&lt;br /&gt;
Problem #171 became a mapping rather than an exclusion.&lt;br /&gt;
The unfiltered source had 463 overlap pairs, a count that includes other issue&lt;br /&gt;
classes and is not the exclusion baseline.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exclude every row participating in&lt;br /&gt;
an otherwise eligible overlap.  Do not select one interval from a pair.  The&lt;br /&gt;
shared projection must define an execution-local source-row identity, include&lt;br /&gt;
this complete query in code, and prove zero overlaps remain before Problem&lt;br /&gt;
#166 sequence closure.&lt;br /&gt;
&lt;br /&gt;
Implemented after independent exclusions as &amp;lt;code&amp;gt;problem_177&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 428 rows are audit&lt;br /&gt;
evidence; sanity proves that no refreshed eligible intervals overlap.  Remove&lt;br /&gt;
only when no otherwise eligible intervals overlap.&lt;br /&gt;
&lt;br /&gt;
== * (#178) ATTENDANCE date components disagree with A_date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The redundant source Day, Mo, and Yr values should describe&lt;br /&gt;
&amp;lt;code&amp;gt;A_date&amp;lt;/code&amp;gt;.  Some day or month values disagree.  Discarding the&lt;br /&gt;
redundant values would hide a source conflict.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT attendance.*&lt;br /&gt;
  FROM clean.attendance AS attendance&lt;br /&gt;
    WHERE attendance.day IS DISTINCT FROM&lt;br /&gt;
        EXTRACT(DAY FROM attendance.a_date)::integer&lt;br /&gt;
      OR attendance.mo IS DISTINCT FROM&lt;br /&gt;
        EXTRACT(MONTH FROM attendance.a_date)::integer&lt;br /&gt;
      OR attendance.yr IS DISTINCT FROM&lt;br /&gt;
        EXTRACT(YEAR FROM attendance.a_date)::integer&lt;br /&gt;
  ORDER BY attendance.a_date,&lt;br /&gt;
           attendance.a_b_animid,&lt;br /&gt;
           attendance.a_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-14 snapshot returned 79 distinct rows.  Day disagreed in 46 rows,&lt;br /&gt;
month in 33, and year in none.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: exactly exclude every returned&lt;br /&gt;
row.  Do not silently prefer one component representation.  Include the query&lt;br /&gt;
beside the Problem #178 predicate and prove all eligible redundant components&lt;br /&gt;
agree with &amp;lt;code&amp;gt;A_date&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Implemented as &amp;lt;code&amp;gt;problem_178&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/attendance_projection.sql&amp;lt;/code&amp;gt;.  The dated 79 rows are audit&lt;br /&gt;
evidence; sanity proves every refreshed eligible redundant date component&lt;br /&gt;
agrees with &amp;lt;code&amp;gt;A_date&amp;lt;/code&amp;gt;.  Remove only when every source row agrees.&lt;br /&gt;
&lt;br /&gt;
Problems #165, #166, #168 through #170, and #172 through #178 overlap and must&lt;br /&gt;
not be summed.  Problem #171 is profiled but mapped, not excluded.  In the&lt;br /&gt;
approved dependency order -- #165, independent #168-#170/#172-#176/#178,&lt;br /&gt;
#177, then #166 -- the refreshed union excluded 13,714 rows and left 120,385&lt;br /&gt;
eligible rows.  Refresh each issue, the overlap set, the final sequence closure,&lt;br /&gt;
and the union before implementation.&lt;br /&gt;
&lt;br /&gt;
On 2026-09-15 the investigator approved replacing complete-row fingerprints&lt;br /&gt;
and frozen issue-membership hashes with semantic validation.  The documented&lt;br /&gt;
predicates and production contracts now determine acceptance.  The conversion&lt;br /&gt;
uses an execution-local &amp;lt;code&amp;gt;source_id&amp;lt;/code&amp;gt; to preserve row multiplicity and&lt;br /&gt;
prove relational parity; it is not a durable source identity.  Historical&lt;br /&gt;
counts remain dated audit evidence and must be refreshed with&lt;br /&gt;
&amp;lt;code&amp;gt;attendance_profile&amp;lt;/code&amp;gt;, including rebuilding the handoff Encounter Log&lt;br /&gt;
when testing a refreshed source.&lt;br /&gt;
&lt;br /&gt;
== (#179) ATTENDANCE observers require standardized PEOPLE matching ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Exact matching would create separate &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; rows for case and&lt;br /&gt;
edge-whitespace variants of the same observer identifier.  This contradicts&lt;br /&gt;
the case-equivalent unique index and the established conversion convention.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Run against the refreshed &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;codes&amp;lt;/code&amp;gt; snapshot after&lt;br /&gt;
food support loading:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH raw_observers AS (&lt;br /&gt;
  SELECT a_observer1_id::text AS raw_person&lt;br /&gt;
    FROM clean.attendance&lt;br /&gt;
    WHERE a_observer1_id IS NOT NULL&lt;br /&gt;
  UNION&lt;br /&gt;
  SELECT a_observer2_id::text&lt;br /&gt;
    FROM clean.attendance&lt;br /&gt;
    WHERE a_observer2_id IS NOT NULL&lt;br /&gt;
), standardized AS (&lt;br /&gt;
  SELECT raw_person,&lt;br /&gt;
         LOWER(NORMALIZE(BTRIM(raw_person))) AS standard_person&lt;br /&gt;
    FROM raw_observers&lt;br /&gt;
)&lt;br /&gt;
SELECT standardized.standard_person,&lt;br /&gt;
       ARRAY_AGG(standardized.raw_person&lt;br /&gt;
                 ORDER BY standardized.raw_person) AS raw_spellings,&lt;br /&gt;
       ARRAY_AGG(people.person ORDER BY people.person)&lt;br /&gt;
         FILTER (WHERE people.person IS NOT NULL) AS existing_codes&lt;br /&gt;
  FROM standardized&lt;br /&gt;
  LEFT JOIN codes.people&lt;br /&gt;
    ON LOWER(NORMALIZE(BTRIM(people.person))) =&lt;br /&gt;
         standardized.standard_person&lt;br /&gt;
  GROUP BY standardized.standard_person&lt;br /&gt;
  ORDER BY standardized.standard_person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The complete refreshed profile contained 932 raw spellings in 718 standardized&lt;br /&gt;
classes.  Sixteen spellings had edge whitespace and 305 raw spellings mapped&lt;br /&gt;
to a different canonical code after trimming and case-equivalent resolution.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14: retain the case-equivalent&lt;br /&gt;
&amp;lt;code&amp;gt;people_person_uniquenocase&amp;lt;/code&amp;gt; index and the edge-whitespace constraint.&lt;br /&gt;
Map every raw attendance spelling through the standardized key defined in&lt;br /&gt;
Problem #167.  Existing &amp;lt;code&amp;gt;PEOPLE.Person&amp;lt;/code&amp;gt; spellings take precedence;&lt;br /&gt;
each previously unseen standardized class receives one deterministic support&lt;br /&gt;
code.  Sanity must prove every raw spelling maps exactly once and every&lt;br /&gt;
standardized class produces exactly one canonical code.  This issue excludes&lt;br /&gt;
no rows except the newly recognized standardized observer equalities under&lt;br /&gt;
Problem #168.&lt;br /&gt;
&lt;br /&gt;
== (#180) load_finish deactivates ATTENDANCE observer codes ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The existing finalization step deactivates &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; and every&lt;br /&gt;
&amp;lt;code&amp;gt;PEOPLE.Person&amp;lt;/code&amp;gt; containing a slash.  After attendance loading this&lt;br /&gt;
makes three referenced observer codes inactive: &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;N/A&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;SELEMANI/YAHAYA&amp;lt;/code&amp;gt;.  They account for 37,655&lt;br /&gt;
Recorder or Observer2 uses, violating the requirement that attendance observer&lt;br /&gt;
codes remain active.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Run after attendance and &amp;lt;code&amp;gt;load_finish&amp;lt;/code&amp;gt;:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT people.person,&lt;br /&gt;
       COUNT(*) AS observer_uses&lt;br /&gt;
  FROM (&lt;br /&gt;
    SELECT arrivals_a.recorder AS person FROM sokwedb.arrivals_a&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT arrivals_a.observer2 FROM sokwedb.arrivals_a&lt;br /&gt;
  ) AS attendance_observers&lt;br /&gt;
    JOIN codes.people USING (person)&lt;br /&gt;
  WHERE NOT people.active&lt;br /&gt;
  GROUP BY people.person&lt;br /&gt;
  ORDER BY people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed query returned three codes and 37,655 observer uses:&lt;br /&gt;
&amp;lt;code&amp;gt;N/A&amp;lt;/code&amp;gt; 54 times, &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; 37,599 times, and&lt;br /&gt;
&amp;lt;code&amp;gt;SELEMANI/YAHAYA&amp;lt;/code&amp;gt; twice.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-14 as part of the active raw-observer&lt;br /&gt;
policy: &amp;lt;code&amp;gt;conversion/load_finish.sql&amp;lt;/code&amp;gt; retains active status for any&lt;br /&gt;
person referenced by &amp;lt;code&amp;gt;ARRIVALS_A.Recorder&amp;lt;/code&amp;gt; or&lt;br /&gt;
&amp;lt;code&amp;gt;ARRIVALS_A.Observer2&amp;lt;/code&amp;gt;.  Problem #39 finalization remains unchanged&lt;br /&gt;
for unreferenced &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;, and slash-containing codes.&lt;br /&gt;
This issue excludes no attendance rows.  Remove the exception only when no&lt;br /&gt;
attendance detail references a code that finalization would deactivate.&lt;br /&gt;
&lt;br /&gt;
== (#181) COLOBUS descriptive text fields contain NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access &amp;lt;code&amp;gt;COLOBUS&amp;lt;/code&amp;gt; source contains NULL in eight descriptive text&lt;br /&gt;
fields whose production destinations are required text columns.  NULL occurs&lt;br /&gt;
in &amp;lt;code&amp;gt;col_hunter_id&amp;lt;/code&amp;gt; 722 times, &amp;lt;code&amp;gt;col_killer_id&amp;lt;/code&amp;gt; 1,438 times,&lt;br /&gt;
&amp;lt;code&amp;gt;col_victim_age&amp;lt;/code&amp;gt; 1,447 times, &amp;lt;code&amp;gt;col_colobus_details&amp;lt;/code&amp;gt; 1,652 times,&lt;br /&gt;
&amp;lt;code&amp;gt;col_veg_details&amp;lt;/code&amp;gt; 2,073 times, &amp;lt;code&amp;gt;col_hunt_failure_details&amp;lt;/code&amp;gt; 1,943 times,&lt;br /&gt;
&amp;lt;code&amp;gt;col_comments&amp;lt;/code&amp;gt; 885 times, and &amp;lt;code&amp;gt;col_hunt_description&amp;lt;/code&amp;gt;&lt;br /&gt;
1,803 times.&lt;br /&gt;
&lt;br /&gt;
The seven colobus-specific values map to &amp;lt;code&amp;gt;COLOBUS.Hunters&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;COLOBUS.Killers&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;COLOBUS.VictimsAges&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;COLOBUS.Details&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;COLOBUS.Vegetation&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;COLOBUS.FailureDetails&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;COLOBUS.Description&amp;lt;/code&amp;gt;.&lt;br /&gt;
Comments map separately to the shared &amp;lt;code&amp;gt;EVENTS.Notes&amp;lt;/code&amp;gt; column.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (WHERE col_hunter_id IS NULL) AS hunters,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_killer_id IS NULL) AS killers,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_victim_age IS NULL) AS victims_ages,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_colobus_details IS NULL) AS troop_details,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_veg_details IS NULL) AS vegetation,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_hunt_failure_details IS NULL)&lt;br /&gt;
         AS failure_details,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_comments IS NULL) AS comments,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_hunt_description IS NULL) AS descriptions&lt;br /&gt;
  FROM clean.colobus;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed query returned, in column order, 722, 1,438, 1,447, 1,652,&lt;br /&gt;
2,073, 1,943, 885, and 1,803 rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-15: retain the existing production&lt;br /&gt;
schema and map each NULL independently to the empty string.  Implement each&lt;br /&gt;
destination as &amp;lt;code&amp;gt;COALESCE(source_value, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt;; preserve every non-NULL&lt;br /&gt;
source value exactly.  The existing required-text and not-only-spaces&lt;br /&gt;
constraints permit the empty string, so this normalization requires no schema&lt;br /&gt;
change.  Do not concatenate comments with the encounter description, and do&lt;br /&gt;
not exclude a source row solely because any of these eight fields is NULL.&lt;br /&gt;
&lt;br /&gt;
Sanity must prove exact field-by-field &amp;lt;code&amp;gt;COALESCE(source_value, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt;&lt;br /&gt;
parity and zero exclusions attributable only to Problem #181.  Remove this&lt;br /&gt;
normalization only when all eight source fields contain no NULL values.&lt;br /&gt;
&lt;br /&gt;
== (#182) MATING_EVENT community codes differ only by case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access &amp;lt;code&amp;gt;MATING_EVENT.M_cl_community_id&amp;lt;/code&amp;gt; column contains the value &amp;lt;code&amp;gt;kk&amp;lt;/code&amp;gt;,&lt;br /&gt;
which differs only by case from the existing production &amp;lt;code&amp;gt;COMM_IDS.CommID&amp;lt;/code&amp;gt; code&lt;br /&gt;
&amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt;.  Production community codes are case-sensitive.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE m_cl_community_id = &amp;#039;kk&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains 181 such rows.  All other mating community&lt;br /&gt;
values are &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;MT&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigators on 2026-09-15: standardize&lt;br /&gt;
&amp;lt;code&amp;gt;m_cl_community_id = &amp;#039;kk&amp;#039;&amp;lt;/code&amp;gt; to the existing &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt; code in the clean schema.  This&lt;br /&gt;
is a case-only spelling normalization and excludes no rows.  Preserve all&lt;br /&gt;
other source community values unchanged, and require the mating sanity check&lt;br /&gt;
to verify that every resulting community exists in &amp;lt;code&amp;gt;COMM_IDS&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#183) MATING_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;EVENTS.Start&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;EVENTS.Stop&amp;lt;/code&amp;gt; record times to minute precision&lt;br /&gt;
and reject nonzero seconds.  The Access &amp;lt;code&amp;gt;MATING_EVENT.M_time&amp;lt;/code&amp;gt; value is initially&lt;br /&gt;
represented as a timestamp, and merely converting it to a PostgreSQL time&lt;br /&gt;
value does not remove seconds.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;MATING_EVENT&amp;quot;&lt;br /&gt;
  WHERE extract(second FROM &amp;quot;M_time&amp;quot;) &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains one row, dated 2008-10-12 with focal GA, male&lt;br /&gt;
WL, female GA, and time 09:35:30.  Problem #4 remains corrected: the refreshed&lt;br /&gt;
source has no non-midnight time component in &amp;lt;code&amp;gt;M_FOL_date&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigators on 2026-09-15: mating event times are recorded&lt;br /&gt;
to minute precision.  During construction of the tidy schema, truncate&lt;br /&gt;
&amp;lt;code&amp;gt;M_time&amp;lt;/code&amp;gt; to the minute before &amp;lt;code&amp;gt;tidy.sql&amp;lt;/code&amp;gt; converts the column to &amp;lt;code&amp;gt;TIME&amp;lt;/code&amp;gt;.  Do not&lt;br /&gt;
round the value.  The mating sanity check must reject the conversion if a&lt;br /&gt;
value containing seconds nevertheless reaches the clean schema.  This issue&lt;br /&gt;
excludes no rows and does not resolve missing or out-of-window mating times.&lt;br /&gt;
&lt;br /&gt;
== * (#184) MATING_EVENT ordinary boolean flags contain question marks ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;MATINGS.Incest&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;Consort&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;Guarding&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;Courting&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;Camp&amp;lt;/code&amp;gt; are&lt;br /&gt;
required Boolean values.  The Access fields use &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt; for true and blank or NULL&lt;br /&gt;
for false, but one &amp;lt;code&amp;gt;M_court_flag&amp;lt;/code&amp;gt; row and two different &amp;lt;code&amp;gt;M_camp_flag&amp;lt;/code&amp;gt; rows&lt;br /&gt;
contain &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt;.  The unknown values cannot be represented by required Booleans.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE BTRIM(m_court_flag) = &amp;#039;?&amp;#039;&lt;br /&gt;
        OR BTRIM(m_camp_flag) = &amp;#039;?&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains three distinct rows.  No row contains &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt; in&lt;br /&gt;
both fields, and none overlaps the Problem #186 exclusion.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigators on 2026-09-15: map &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt; to true and blank or NULL&lt;br /&gt;
to false for all five fields.  Exclude the three rows containing &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt;.  Sanity&lt;br /&gt;
must reject any other domain value.  This issue independently excludes three&lt;br /&gt;
rows and contributes three rows to the exclusion union.&lt;br /&gt;
&lt;br /&gt;
== (#185) MATING_EVENT records two interference facts in one target field ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;M_interference&amp;lt;/code&amp;gt; records a description or chimp ID.&lt;br /&gt;
&amp;lt;code&amp;gt;M_interference_too_late_flag&amp;lt;/code&amp;gt; records the ID of a chimp who interfered too&lt;br /&gt;
late.  Production has only the single required text column&lt;br /&gt;
&amp;lt;code&amp;gt;MATINGS.Interference&amp;lt;/code&amp;gt;, so copying or concatenating the fields without labels&lt;br /&gt;
would lose which fact each value represents.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (&lt;br /&gt;
         WHERE NULLIF(BTRIM(m_interference), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
               AND NULLIF(BTRIM(m_interference_too_late_flag), &amp;#039;&amp;#039;) IS NOT NULL)&lt;br /&gt;
         AS both_fields,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE NULLIF(BTRIM(m_interference), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
               AND NULLIF(BTRIM(m_interference_too_late_flag), &amp;#039;&amp;#039;) IS NULL)&lt;br /&gt;
         AS interference_only,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE NULLIF(BTRIM(m_interference), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
               AND NULLIF(BTRIM(m_interference_too_late_flag), &amp;#039;&amp;#039;) IS NOT NULL)&lt;br /&gt;
         AS too_late_only&lt;br /&gt;
  FROM clean.mating_event;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains 82 rows with both fields, 3,760 with only&lt;br /&gt;
&amp;lt;code&amp;gt;m_interference&amp;lt;/code&amp;gt;, and 135 with only the too-late field.  No source&lt;br /&gt;
&amp;lt;code&amp;gt;m_interference&amp;lt;/code&amp;gt; contains the labels &amp;lt;code&amp;gt;Interference:&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;Too late:&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: trim both fields and use this&lt;br /&gt;
explicit encoding in &amp;lt;code&amp;gt;MATINGS.Interference&amp;lt;/code&amp;gt;:&lt;br /&gt;
&lt;br /&gt;
* neither field: the empty string;&lt;br /&gt;
* &amp;lt;code&amp;gt;m_interference&amp;lt;/code&amp;gt; only: the trimmed source value;&lt;br /&gt;
* too-late only: &amp;lt;code&amp;gt;Too late: &amp;lt;/code&amp;gt; followed by the trimmed source value; and&lt;br /&gt;
* both fields: &amp;lt;code&amp;gt;Interference: &amp;lt;/code&amp;gt; followed by the trimmed interference value,&lt;br /&gt;
  then &amp;lt;code&amp;gt;; Too late: &amp;lt;/code&amp;gt; followed by the trimmed too-late value.&lt;br /&gt;
&lt;br /&gt;
This preserves the two source facts without changing the production schema.&lt;br /&gt;
Problem #186 separately excludes too-late &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt; values before this encoding.&lt;br /&gt;
This issue excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== * (#186) MATING_EVENT too-late interference contains X instead of an ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The investigators identified &amp;lt;code&amp;gt;M_interference_too_late_flag&amp;lt;/code&amp;gt; as the ID of a&lt;br /&gt;
chimp who interfered too late, but 110 source rows contain &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;.  &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt; is not a&lt;br /&gt;
chimp ID, and the source does not identify an individual who can replace it.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE BTRIM(m_interference_too_late_flag) = &amp;#039;X&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 110 rows span 1976 through 2015 and all use source &amp;lt;code&amp;gt;B-REC&amp;lt;/code&amp;gt;.  Eighty-two&lt;br /&gt;
also contain a separate interference description or ID; 28 do not.  In&lt;br /&gt;
contrast, the 107 chimp-like values occur only in 2012 through 2015 and always&lt;br /&gt;
have an empty main interference field.  The evidence is consistent with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;&lt;br /&gt;
being an older marker, but it does not recover the required chimp ID.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: exclude all 110 rows containing&lt;br /&gt;
&amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt; in &amp;lt;code&amp;gt;m_interference_too_late_flag&amp;lt;/code&amp;gt;.  Do not interpret &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt; as an individual&lt;br /&gt;
or silently discard the too-late fact.  This issue independently excludes 110&lt;br /&gt;
rows, overlaps none of the three Problem #184 rows, and increases the&lt;br /&gt;
exclusion union from 3 to 113 rows.&lt;br /&gt;
&lt;br /&gt;
== * (#187) MATING_EVENT fail flags contain missing and unknown values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;MATINGS.Fail&amp;lt;/code&amp;gt; is a required Boolean.  The investigators confirmed&lt;br /&gt;
that &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;Y&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;y&amp;lt;/code&amp;gt; mean the mating was not complete and map to true, while&lt;br /&gt;
&amp;lt;code&amp;gt;N&amp;lt;/code&amp;gt; maps to false.  The Access source also contains blank, NULL, and &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt;&lt;br /&gt;
values, which require an explicit disposition.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(m_fail_flag), &amp;#039;&amp;#039;), &amp;#039;[blank/NULL]&amp;#039;) AS fail_value,&lt;br /&gt;
       COUNT(*)&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  GROUP BY COALESCE(NULLIF(BTRIM(m_fail_flag), &amp;#039;&amp;#039;), &amp;#039;[blank/NULL]&amp;#039;)&lt;br /&gt;
  ORDER BY fail_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains 21,786 blank or NULL values and 552 &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt; values.&lt;br /&gt;
After Problems #184 and #186, 21,676 blank or NULL values and 551 &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt; values&lt;br /&gt;
remain eligible.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: map &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;Y&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;y&amp;lt;/code&amp;gt; to true;&lt;br /&gt;
map &amp;lt;code&amp;gt;N&amp;lt;/code&amp;gt;, blank, and NULL to false; and exclude rows containing &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt;.  Sanity&lt;br /&gt;
must reject any other domain value.  The &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt; condition independently excludes&lt;br /&gt;
552 rows, overlaps one Problem #184 row and no Problem #186 rows, and increases&lt;br /&gt;
the exclusion union from 113 to 664 rows.  It contributes 551 newly excluded&lt;br /&gt;
rows.&lt;br /&gt;
&lt;br /&gt;
== * (#188) MATING_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the separately governed &amp;lt;code&amp;gt;sdb_no_time&amp;lt;/code&amp;gt; sentinel, production&lt;br /&gt;
&amp;lt;code&amp;gt;EVENTS.Start&amp;lt;/code&amp;gt; cannot be before 04:00 and &amp;lt;code&amp;gt;EVENTS.Stop&amp;lt;/code&amp;gt; cannot be after 20:00.&lt;br /&gt;
Some non-NULL Access mating times fall outside that interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE m_time IS NOT NULL&lt;br /&gt;
        AND (m_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME OR m_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
  ORDER BY m_fol_date, m_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains 25 rows: 24 between 01:00 and 03:48 and one at&lt;br /&gt;
21:38.  None overlaps Problems #184, #186, or #187.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: exclude the 25 rows with non-NULL&lt;br /&gt;
times outside 04:00 through 20:00.  Do not clamp, wrap, or reinterpret their&lt;br /&gt;
times.  This issue independently and newly excludes 25 rows, increasing the&lt;br /&gt;
exclusion union from 664 to 689 rows.&lt;br /&gt;
&lt;br /&gt;
== (#189) MATING_EVENT rows may have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access source contains 34 mating rows with NULL &amp;lt;code&amp;gt;m_time&amp;lt;/code&amp;gt;.  Production&lt;br /&gt;
&amp;lt;code&amp;gt;EVENTS.Start&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;EVENTS.Stop&amp;lt;/code&amp;gt; are required.  Before this decision, only&lt;br /&gt;
aggression, B-record-note, and pantgrunt events could use the established&lt;br /&gt;
&amp;lt;code&amp;gt;sdb_no_time&amp;lt;/code&amp;gt; value to distinguish an unrecorded time from an observed time.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE m_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains 34 rows.  One overlaps the prior exclusion&lt;br /&gt;
union, so excluding missing times would discard 33 additional otherwise&lt;br /&gt;
eligible mating records.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: permit mating events to use the&lt;br /&gt;
existing &amp;lt;code&amp;gt;sdb_no_time&amp;lt;/code&amp;gt; sentinel.  Add &amp;lt;code&amp;gt;sdb_mating_event&amp;lt;/code&amp;gt; to the two owning&lt;br /&gt;
&amp;lt;code&amp;gt;EVENTS&amp;lt;/code&amp;gt; constraint allowlists and document the expanded behavior set.  At the&lt;br /&gt;
production load boundary, map only SQL NULL &amp;lt;code&amp;gt;m_time&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;sdb_no_time&amp;lt;/code&amp;gt;&lt;br /&gt;
for both &amp;lt;code&amp;gt;EVENTS.Start&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;EVENTS.Stop&amp;lt;/code&amp;gt;; preserve every non-NULL source time.&lt;br /&gt;
The existing point-event and paired-sentinel constraints remain in force.&lt;br /&gt;
This issue excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#190) MATING_EVENT swelling values may be missing ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;MATINGS.Swelling&amp;lt;/code&amp;gt; is required, but the Access source contains 572&lt;br /&gt;
NULL values and 16 &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt; values.  The source context for the question marks&lt;br /&gt;
describes unavailable information, including &amp;quot;no estrous state noted&amp;quot; and&lt;br /&gt;
&amp;quot;swelling not recorded&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE m_fem_swelling IS NULL&lt;br /&gt;
        OR BTRIM(m_fem_swelling) = &amp;#039;?&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After prior exclusions, 568 NULL values and all 16 question marks remain&lt;br /&gt;
eligible.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: map both NULL and &amp;lt;code&amp;gt;?&amp;lt;/code&amp;gt; to the&lt;br /&gt;
existing &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; code &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, documented as &amp;quot;Missing data&amp;quot;.  &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; has&lt;br /&gt;
NULL &amp;lt;code&amp;gt;AsNum&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;SSRank&amp;lt;/code&amp;gt;, so these records do not contribute a measured value&lt;br /&gt;
to daily swelling minimum/maximum calculations.  The original missing-value&lt;br /&gt;
encoding remains available in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt;.  This issue excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#191) MATING_EVENT records one-third swelling ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Nine otherwise eligible Access mating rows contain the swelling value &amp;lt;code&amp;gt;0.33&amp;lt;/code&amp;gt;,&lt;br /&gt;
which is absent from &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt;.  Inserting the value into &amp;lt;code&amp;gt;MATINGS&amp;lt;/code&amp;gt;&lt;br /&gt;
reproduces a &amp;lt;code&amp;gt;matings_swelling_fkey&amp;lt;/code&amp;gt; violation.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE BTRIM(m_fem_swelling) = &amp;#039;0.33&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: add &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; code &amp;lt;code&amp;gt;0.33&amp;lt;/code&amp;gt;&lt;br /&gt;
with &amp;lt;code&amp;gt;AsNum&amp;lt;/code&amp;gt; 0.33, &amp;lt;code&amp;gt;SSRank&amp;lt;/code&amp;gt; 3, and description &amp;quot;1/3 swollen&amp;quot;.  Shift the&lt;br /&gt;
existing &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; support rows to ranks 4 through 7 so the&lt;br /&gt;
unique integer rank continues to order measured swelling values correctly.&lt;br /&gt;
Map source &amp;lt;code&amp;gt;0.33&amp;lt;/code&amp;gt; to the new exact code.  This issue excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#192) MATING_EVENT extractors require PEOPLE support ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;MATINGS.ExtractedBy&amp;lt;/code&amp;gt; is required and must reference an active&lt;br /&gt;
&amp;lt;code&amp;gt;PEOPLE.Person&amp;lt;/code&amp;gt;.  The source contains 20,747 NULL extractor values, one&lt;br /&gt;
case-only spelling difference, and three populated names absent from the&lt;br /&gt;
conversion &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT m_extracted_by, COUNT(*)&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE NULLIF(BTRIM(m_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
        OR NOT EXISTS (&lt;br /&gt;
             SELECT 1&lt;br /&gt;
               FROM clean.people&lt;br /&gt;
               WHERE LOWER(NORMALIZE(BTRIM(people.person))) =&lt;br /&gt;
                     LOWER(NORMALIZE(BTRIM(mating_event.m_extracted_by))))&lt;br /&gt;
  GROUP BY m_extracted_by&lt;br /&gt;
  ORDER BY m_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;Karen McLellan&amp;lt;/code&amp;gt; matches the existing stored spelling &amp;lt;code&amp;gt;KAREN MCLELLAN&amp;lt;/code&amp;gt;.&lt;br /&gt;
&amp;lt;code&amp;gt;Deus Mjungu&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FELDBLUM&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;Joclyn Antonio&amp;lt;/code&amp;gt; have no standardized match.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: map a NULL, empty, or&lt;br /&gt;
whitespace-only extractor to &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;; trim populated values; reuse the exact&lt;br /&gt;
stored &amp;lt;code&amp;gt;PEOPLE.Person&amp;lt;/code&amp;gt; spelling for standardized case-equivalent matches; and&lt;br /&gt;
add &amp;lt;code&amp;gt;Deus Mjungu&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FELDBLUM&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;Joclyn Antonio&amp;lt;/code&amp;gt; as active support people.&lt;br /&gt;
Use &amp;lt;code&amp;gt;LOWER(NORMALIZE(BTRIM(value)))&amp;lt;/code&amp;gt; as the deterministic matching key.  This&lt;br /&gt;
issue excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#193) MATING_EVENT comments require a production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access source stores mating comments separately from the full mating&lt;br /&gt;
description.  Production has one shared event-notes field, &amp;lt;code&amp;gt;EVENTS.Notes&amp;lt;/code&amp;gt;,&lt;br /&gt;
and silently concatenating the two source fields would erase that distinction.&lt;br /&gt;
The source contains 14,294 NULL comments and 15,134 comments that are NULL or&lt;br /&gt;
empty after trimming.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (WHERE m_comments IS NULL) AS null_comments,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE NULLIF(BTRIM(m_comments), &amp;#039;&amp;#039;) IS NULL) AS blank_comments&lt;br /&gt;
  FROM clean.mating_event;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: map &amp;lt;code&amp;gt;m_comments&amp;lt;/code&amp;gt; alone to&lt;br /&gt;
&amp;lt;code&amp;gt;EVENTS.Notes&amp;lt;/code&amp;gt;.  Map SQL NULL to the empty string with&lt;br /&gt;
&amp;lt;code&amp;gt;COALESCE(m_comments, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt; and preserve every non-NULL source value exactly.&lt;br /&gt;
Do not concatenate &amp;lt;code&amp;gt;m_full_description&amp;lt;/code&amp;gt; into event notes.  This issue excludes&lt;br /&gt;
no rows.&lt;br /&gt;
&lt;br /&gt;
== (#194) MATING_EVENT full descriptions require a mating-specific destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access source stores a full mating description in &amp;lt;code&amp;gt;m_full_description&amp;lt;/code&amp;gt;.&lt;br /&gt;
&amp;lt;code&amp;gt;BRECORD_NOTES&amp;lt;/code&amp;gt; cannot preserve it because that detail table can attach only&lt;br /&gt;
to a B-record event, cannot share the mating EID, and requires unrelated&lt;br /&gt;
B-record text fields.  The source contains 18,212 descriptions that are NULL&lt;br /&gt;
or empty after trimming.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (WHERE m_full_description IS NULL)&lt;br /&gt;
         AS null_descriptions,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE NULLIF(BTRIM(m_full_description), &amp;#039;&amp;#039;) IS NULL)&lt;br /&gt;
         AS blank_descriptions&lt;br /&gt;
  FROM clean.mating_event;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: add the required text column&lt;br /&gt;
&amp;lt;code&amp;gt;MATINGS.Description&amp;lt;/code&amp;gt; and map &amp;lt;code&amp;gt;m_full_description&amp;lt;/code&amp;gt; to it.  Map SQL NULL to the&lt;br /&gt;
empty string with &amp;lt;code&amp;gt;COALESCE(m_full_description, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt; and preserve every&lt;br /&gt;
non-NULL source value exactly.  The production constraint permits the empty&lt;br /&gt;
string but rejects nonempty whitespace-only values.&lt;br /&gt;
&lt;br /&gt;
The permanent schema and documentation change was implemented as append-only&lt;br /&gt;
master commit &amp;lt;code&amp;gt;a628fdd&amp;lt;/code&amp;gt; (&amp;lt;code&amp;gt;feat(matings): preserve mating descriptions&amp;lt;/code&amp;gt;) and&lt;br /&gt;
fast-forwarded to local, origin, private, and blessed master.  This issue&lt;br /&gt;
excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#195) MATING_EVENT sources require MATING_SOURCES support ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;MATINGS.Source&amp;lt;/code&amp;gt; is required and references&lt;br /&gt;
&amp;lt;code&amp;gt;MATING_SOURCES.Source&amp;lt;/code&amp;gt;, but the source code table is empty.  The Access source&lt;br /&gt;
contains mixed-case source labels while the production support key requires&lt;br /&gt;
uppercase values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT m_source, COUNT(*) AS row_count,&lt;br /&gt;
       MIN(m_fol_date) AS first_date,&lt;br /&gt;
       MAX(m_fol_date) AS last_date&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  GROUP BY m_source&lt;br /&gt;
  ORDER BY m_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed source contains &amp;lt;code&amp;gt;B-REC&amp;lt;/code&amp;gt; 26,066 times over 1976 through 2018,&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt; 393 times over 2000 through 2007, and &amp;lt;code&amp;gt;Bobu&amp;lt;/code&amp;gt; 24 times in 2000.  The&lt;br /&gt;
repository provides no more authoritative definition of &amp;lt;code&amp;gt;Bobu&amp;lt;/code&amp;gt; than its&lt;br /&gt;
source label.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: add literal support codes &amp;lt;code&amp;gt;B-REC&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;TIKI&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;BOBU&amp;lt;/code&amp;gt;, described respectively as &amp;quot;Mating record sourced from&lt;br /&gt;
B-REC&amp;quot;, &amp;quot;Mating record sourced from Tiki&amp;quot;, and &amp;quot;Mating record sourced from&lt;br /&gt;
Bobu&amp;quot;.  At the production boundary map source &amp;lt;code&amp;gt;B-REC&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;B-REC&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;TIKI&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;Bobu&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;BOBU&amp;lt;/code&amp;gt;.  Do not claim semantics beyond the source labels.&lt;br /&gt;
Sanity must reject any refreshed source value outside this approved domain and&lt;br /&gt;
verify that every transformed code exists exactly once.  This issue excludes&lt;br /&gt;
no rows.&lt;br /&gt;
&lt;br /&gt;
== * (#196) MATING_EVENT participants are absent from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;ROLES.Participant&amp;lt;/code&amp;gt; is required and references&lt;br /&gt;
&amp;lt;code&amp;gt;BIOGRAPHY_DATA.AnimID&amp;lt;/code&amp;gt;.  The first exclusion-free loader run failed while&lt;br /&gt;
inserting female participant &amp;lt;code&amp;gt;STR&amp;lt;/code&amp;gt; for conversion identity 117.  The source&lt;br /&gt;
also contains other participant values that are not production biography&lt;br /&gt;
identifiers, including category labels, unknown markers, blanks, and&lt;br /&gt;
unresolved spellings.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH eligible AS (&lt;br /&gt;
  SELECT *&lt;br /&gt;
    FROM clean.mating_event&lt;br /&gt;
    WHERE COALESCE(BTRIM(m_court_flag), &amp;#039;&amp;#039;) &amp;lt;&amp;gt; &amp;#039;?&amp;#039;&lt;br /&gt;
          AND COALESCE(BTRIM(m_camp_flag), &amp;#039;&amp;#039;) &amp;lt;&amp;gt; &amp;#039;?&amp;#039;&lt;br /&gt;
          AND COALESCE(BTRIM(m_interference_too_late_flag), &amp;#039;&amp;#039;) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
          AND COALESCE(BTRIM(m_fail_flag), &amp;#039;&amp;#039;) &amp;lt;&amp;gt; &amp;#039;?&amp;#039;&lt;br /&gt;
          AND (m_time IS NULL&lt;br /&gt;
               OR m_time BETWEEN &amp;#039;04:00&amp;#039;::TIME AND &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
)&lt;br /&gt;
SELECT DISTINCT eligible.m_conversion_id&lt;br /&gt;
  FROM eligible&lt;br /&gt;
    CROSS JOIN LATERAL (&lt;br /&gt;
      VALUES (eligible.m_male), (eligible.m_female)&lt;br /&gt;
    ) AS participant(animid)&lt;br /&gt;
  WHERE NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
      FROM sokwedb.biography_data&lt;br /&gt;
      WHERE biography_data.animid = BTRIM(participant.animid))&lt;br /&gt;
  ORDER BY eligible.m_conversion_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The refreshed post-follow database contains 574 missing participant&lt;br /&gt;
occurrences in 566 otherwise eligible source rows, representing 34 distinct&lt;br /&gt;
trimmed values.  The first failure is identity 117, dated 2000-10-14, with&lt;br /&gt;
focal &amp;lt;code&amp;gt;BE&amp;lt;/code&amp;gt;, male &amp;lt;code&amp;gt;SL&amp;lt;/code&amp;gt;, and female &amp;lt;code&amp;gt;STR&amp;lt;/code&amp;gt;.  The failed first chunk rolled back&lt;br /&gt;
and left zero &amp;lt;code&amp;gt;MATE&amp;lt;/code&amp;gt; events.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: exclude every source row for which&lt;br /&gt;
either trimmed participant is absent from &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt;.  Do not create&lt;br /&gt;
biography rows or substitute participant identities without source evidence.&lt;br /&gt;
Sanity must verify that this condition selects exactly 566 distinct eligible&lt;br /&gt;
source identities in the retained snapshot.  The rows overlap none of the&lt;br /&gt;
prior 689-row exclusion union and increase it to 1,255 rows, leaving 25,228&lt;br /&gt;
rows eligible before later failures.&lt;br /&gt;
&lt;br /&gt;
== * (#197) MATING_EVENT participants fall outside biography study dates ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After excluding Problem #196 rows, the loader failed when actor &amp;lt;code&amp;gt;FAM&amp;lt;/code&amp;gt;&lt;br /&gt;
participated on 2014-09-10, after the individual&amp;#039;s 2012-04-26 departure date.&lt;br /&gt;
Production requires every role participant to be biography-valid on the event&lt;br /&gt;
date.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT DISTINCT mating_event.m_conversion_id&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
    CROSS JOIN LATERAL (&lt;br /&gt;
      VALUES (mating_event.m_male), (mating_event.m_female)&lt;br /&gt;
    ) AS participant(animid)&lt;br /&gt;
    JOIN sokwedb.biography_data&lt;br /&gt;
      ON biography_data.animid = BTRIM(participant.animid)&lt;br /&gt;
  WHERE mating_event.m_fol_date &amp;lt; biography_data.entrydate&lt;br /&gt;
        OR mating_event.m_fol_date &amp;gt; biography_data.departdate&lt;br /&gt;
  ORDER BY mating_event.m_conversion_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After applying Problems #184, #186, #187, #188, and #196, 28 participant&lt;br /&gt;
occurrences in 28 source rows fall outside study dates.  Three are male actors&lt;br /&gt;
after departure, 13 are female actees after departure, and 12 are female&lt;br /&gt;
actees before entry.  The affected participants are &amp;lt;code&amp;gt;AT&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;CF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FAM&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OR&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;SF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;SH&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;SIF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;SP&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;WN&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YD&amp;lt;/code&amp;gt;.  The failed first chunk rolled back&lt;br /&gt;
and left zero &amp;lt;code&amp;gt;MATE&amp;lt;/code&amp;gt; events.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: exclude all 28 rows.  Do not alter&lt;br /&gt;
source dates or biography study spans.  Sanity must verify the exact 28-row&lt;br /&gt;
condition after the prior exclusions.  This increases the exclusion union&lt;br /&gt;
from 1,255 to 1,283 rows and leaves 25,200 rows eligible before later&lt;br /&gt;
failures.&lt;br /&gt;
&lt;br /&gt;
== * (#198) MATING_EVENT focal values are absent from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After Problems #196 and #197, the loader failed while creating an Other watch&lt;br /&gt;
for focal &amp;lt;code&amp;gt;LUT&amp;lt;/code&amp;gt;, which is absent from &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt;.  Production requires&lt;br /&gt;
every &amp;lt;code&amp;gt;WATCHES.AnimID&amp;lt;/code&amp;gt;, including mating-only Other watches, to be a biography&lt;br /&gt;
identifier or the established no-focal sentinel.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(m_fol_b_focal_animid), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS focal,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
  WHERE NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
      FROM sokwedb.biography_data&lt;br /&gt;
      WHERE biography_data.animid = COALESCE(&lt;br /&gt;
              NULLIF(BTRIM(mating_event.m_fol_b_focal_animid), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;))&lt;br /&gt;
  GROUP BY COALESCE(NULLIF(BTRIM(m_fol_b_focal_animid), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;)&lt;br /&gt;
  ORDER BY focal;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After the prior exclusions, 311 otherwise eligible rows use 51 unsupported&lt;br /&gt;
focal values.  Three source rows have unique case-insensitive biography&lt;br /&gt;
matches: &amp;lt;code&amp;gt;Gb&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;GB&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;lam&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;LAM&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;sif&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;SIF&amp;lt;/code&amp;gt;.  The remaining 308&lt;br /&gt;
rows use 48 unsupported values, many representing groups, multiple animals,&lt;br /&gt;
or missing-target markers.  The failed first chunk rolled back and left zero&lt;br /&gt;
&amp;lt;code&amp;gt;MATE&amp;lt;/code&amp;gt; events.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: normalize only the three proven&lt;br /&gt;
case variants in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt;, then exclude the 308 rows whose normalized focal is&lt;br /&gt;
still absent from &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt;.  Do not infer individual focal identities&lt;br /&gt;
from group or multi-animal values.  Sanity must verify the exact 308-row&lt;br /&gt;
residual after prior exclusions.  This increases the exclusion union from&lt;br /&gt;
1,283 to 1,591 rows and leaves 24,892 rows eligible before later failures.&lt;br /&gt;
&lt;br /&gt;
== * (#199) MATING_EVENT focal dates are outside BIOGRAPHY_DATA study dates ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After Problem #198, the loader failed while creating an Other watch for focal&lt;br /&gt;
&amp;lt;code&amp;gt;POR&amp;lt;/code&amp;gt; on 2016-09-22.  &amp;lt;code&amp;gt;POR&amp;lt;/code&amp;gt; departed on 2014-03-13, and production requires a&lt;br /&gt;
watch date to fall within the focal animal&amp;#039;s biography study dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT mating_event.m_conversion_id,&lt;br /&gt;
       mating_event.m_fol_date,&lt;br /&gt;
       focal_biography.animid,&lt;br /&gt;
       focal_biography.entrydate,&lt;br /&gt;
       focal_biography.departdate&lt;br /&gt;
  FROM clean.mating_event&lt;br /&gt;
    JOIN sokwedb.biography_data AS focal_biography&lt;br /&gt;
      ON focal_biography.animid = COALESCE(&lt;br /&gt;
           NULLIF(BTRIM(mating_event.m_fol_b_focal_animid), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;)&lt;br /&gt;
  WHERE mating_event.m_fol_date &amp;lt; focal_biography.entrydate&lt;br /&gt;
        OR mating_event.m_fol_date &amp;gt; focal_biography.departdate&lt;br /&gt;
  ORDER BY mating_event.m_conversion_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After the prior exclusions, 43 otherwise eligible rows have focal dates&lt;br /&gt;
outside the focal&amp;#039;s study dates.  Forty-one rows use &amp;lt;code&amp;gt;POR&amp;lt;/code&amp;gt; after departure;&lt;br /&gt;
two use &amp;lt;code&amp;gt;SN&amp;lt;/code&amp;gt; before entry.  The failure occurred after 16,000 rows had&lt;br /&gt;
committed in earlier chunks, while the failing chunk rolled back.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-16: exclude all 43 rows.  Do not alter&lt;br /&gt;
source dates or biography study spans.  Sanity must verify the exact 43-row&lt;br /&gt;
condition after the prior exclusions.  This increases the exclusion union&lt;br /&gt;
from 1,591 to 1,634 rows and leaves 24,849 rows eligible before later&lt;br /&gt;
failures.  Rebuild the post-follow baseline before retrying so committed MATE&lt;br /&gt;
rows and any mating-created Other watches are removed together.&lt;br /&gt;
&lt;br /&gt;
== (#200) COLOBUS encounters have no recorded Stop ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 204 COLOBUS source rows with a recorded Start but no Stop.  The&lt;br /&gt;
general EVENTS sentinel rules did not permit preserving a known Start with an&lt;br /&gt;
unknown Stop for behavior &amp;lt;code&amp;gt;COL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT col_encounter_date, col_fol_b_focal_chimp_id, col_begin_time&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_end_time IS NULL&lt;br /&gt;
  ORDER BY col_encounter_date, col_fol_b_focal_chimp_id, col_begin_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: retain each recorded Start and map only the&lt;br /&gt;
missing Stop to &amp;lt;code&amp;gt;sdb_no_time&amp;lt;/code&amp;gt;.  Add a narrow, documented EVENTS exception for&lt;br /&gt;
COL with a non-sentinel Start and sentinel Stop.  Do not exclude these rows or&lt;br /&gt;
allow a sentinel COL Start.  Validation must find exactly 204 converted COL&lt;br /&gt;
events with the sentinel Stop.&lt;br /&gt;
&lt;br /&gt;
== * (#201) COLOBUS encounter times violate the approved time range ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Fifteen rows cannot be represented without changing a recorded event time:&lt;br /&gt;
eight have Start after Stop, two have Start after 20:00, and seven have Stop&lt;br /&gt;
after 20:00.  The two late Starts overlap the Start-after-Stop class, so the&lt;br /&gt;
distinct union is 15 rows.  None has a missing Stop.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT col_encounter_date, col_fol_b_focal_chimp_id,&lt;br /&gt;
       col_begin_time, col_end_time&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_begin_time &amp;gt; col_end_time&lt;br /&gt;
        OR col_begin_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME&lt;br /&gt;
        OR col_end_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME&lt;br /&gt;
  ORDER BY col_encounter_date, col_fol_b_focal_chimp_id, col_begin_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: exclude exactly this 15-row union at the loader&lt;br /&gt;
boundary and preserve all other recorded times.  Sanity must fail if the&lt;br /&gt;
predicate no longer identifies exactly 15 rows.  This establishes an&lt;br /&gt;
exclusion union of 15 rows before Problem #210.&lt;br /&gt;
&lt;br /&gt;
== (#202) COLOBUS StartMap is missing from seven rows ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Seven source rows have no &amp;lt;code&amp;gt;COL_begin_time_MAP&amp;lt;/code&amp;gt;.  Of the 2,240 recorded values,&lt;br /&gt;
126 intentionally differ from the standard 15-minute calculation and must not&lt;br /&gt;
be recalculated.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (WHERE col_begin_time_map IS NULL) AS missing,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE col_begin_time_map IS NOT NULL&lt;br /&gt;
               AND col_begin_time_map IS DISTINCT FROM&lt;br /&gt;
                 DATE_BIN(&amp;#039;15 minutes&amp;#039;,&lt;br /&gt;
                          DATE &amp;#039;1960-07-04&amp;#039; + col_begin_time&lt;br /&gt;
                            + INTERVAL &amp;#039;8 minutes&amp;#039;,&lt;br /&gt;
                          TIMESTAMP &amp;#039;1960-07-04 00:00&amp;#039;)::TIME) AS adjusted&lt;br /&gt;
  FROM clean.colobus;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: derive only the seven missing values at the&lt;br /&gt;
production load boundary with the displayed &amp;lt;code&amp;gt;DATE_BIN&amp;lt;/code&amp;gt; expression.  Preserve&lt;br /&gt;
all recorded values exactly.  This excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#203) COLOBUS Hunt, HuntCert, and Kill can be unknown ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The source contains 15 NULL Hunt flags, 16 NULL HuntCert flags, and 221 NULL&lt;br /&gt;
Kill flags, but the corresponding production columns were &amp;lt;code&amp;gt;NOT NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (WHERE col_hunt_flag IS NULL) AS hunt_unknown,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_hunt_cert IS NULL) AS huntcert_unknown,&lt;br /&gt;
       COUNT(*) FILTER (WHERE col_kill_flag IS NULL) AS kill_unknown&lt;br /&gt;
  FROM clean.colobus;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: make all three production columns nullable,&lt;br /&gt;
map &amp;lt;code&amp;gt;Y&amp;lt;/code&amp;gt; to true and &amp;lt;code&amp;gt;N&amp;lt;/code&amp;gt; to false, and preserve source NULL as SQL NULL.  This&lt;br /&gt;
excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#204) COLOBUS FocalHuntCert can be unknown ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 744 NULL focal-hunt-certainty values.  Documentation described an&lt;br /&gt;
unknown state, but the production column was &amp;lt;code&amp;gt;NOT NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*)&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_focal_hunt_cert IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: resolve the schema/documentation conflict in&lt;br /&gt;
favor of the documented unknown state.  Make &amp;lt;code&amp;gt;FocalHuntCert&amp;lt;/code&amp;gt; nullable, map&lt;br /&gt;
&amp;lt;code&amp;gt;Y&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;N&amp;lt;/code&amp;gt; to booleans, and preserve all source NULLs.  This excludes no&lt;br /&gt;
rows.&lt;br /&gt;
&lt;br /&gt;
== (#205) COLOBUS NumKills is absent for several distinct states ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 389 rows without a total kill count: seven have &amp;lt;code&amp;gt;Kill = &amp;#039;Y&amp;#039;&amp;lt;/code&amp;gt;, 163&lt;br /&gt;
have &amp;lt;code&amp;gt;Kill = &amp;#039;N&amp;#039;&amp;lt;/code&amp;gt;, and 219 also have an unknown Kill flag.  Seven meaningful&lt;br /&gt;
recorded male totals exceed the recorded overall total and must not be&lt;br /&gt;
rewritten.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT col_kill_flag, COUNT(*)&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_num_kills_total IS NULL&lt;br /&gt;
  GROUP BY col_kill_flag&lt;br /&gt;
  ORDER BY col_kill_flag NULLS LAST;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: make &amp;lt;code&amp;gt;NumKills&amp;lt;/code&amp;gt; nullable.  Preserve every&lt;br /&gt;
recorded total; for a missing total map &amp;lt;code&amp;gt;Kill = &amp;#039;Y&amp;#039;&amp;lt;/code&amp;gt; to 1, &amp;lt;code&amp;gt;Kill = &amp;#039;N&amp;#039;&amp;lt;/code&amp;gt; to 0,&lt;br /&gt;
and unknown Kill to SQL NULL.  Preserve the seven male/total differences and&lt;br /&gt;
verify source-to-target parity.  This excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#206) COLOBUS FemaleNumKills has no supported derivation ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The source does not provide an independently supported total female kill&lt;br /&gt;
count.  Deriving it from other totals would invent data and can be invalid for&lt;br /&gt;
the seven rows where male kills exceed the overall total.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*)&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_num_kills_total IS NOT NULL&lt;br /&gt;
        AND col_num_kills_by_males &amp;gt; col_num_kills_total;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The query returns seven rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: make &amp;lt;code&amp;gt;FemaleNumKills&amp;lt;/code&amp;gt; nullable and always load&lt;br /&gt;
SQL NULL.  Do not derive it by subtraction.  This excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#207) COLOBUS focal kill counts contain fractions ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Three focal-male and three focal-female source counts have fractional values.&lt;br /&gt;
The production &amp;lt;code&amp;gt;INTEGER&amp;lt;/code&amp;gt; columns cannot preserve them.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (&lt;br /&gt;
         WHERE col_num_kills_by_focal_male &amp;lt;&amp;gt;&lt;br /&gt;
               TRUNC(col_num_kills_by_focal_male)) AS focal_male,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE col_num_kills_by_focal_female &amp;lt;&amp;gt;&lt;br /&gt;
               TRUNC(col_num_kills_by_focal_female)) AS focal_female&lt;br /&gt;
  FROM clean.colobus;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: change &amp;lt;code&amp;gt;FocalMaleNumKills&amp;lt;/code&amp;gt; and&lt;br /&gt;
&amp;lt;code&amp;gt;FocalFemaleNumKills&amp;lt;/code&amp;gt; to exact &amp;lt;code&amp;gt;NUMERIC&amp;lt;/code&amp;gt; columns and preserve each source&lt;br /&gt;
value.  This excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#208) COLOBUS group-size codes are unsupported and can be absent ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The support table was empty.  Source values use codes &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt;, and 611&lt;br /&gt;
rows have no group-size code.  Missing values must not be changed to code &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt;,&lt;br /&gt;
which has a distinct recorded meaning.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT col_colobus_group_size, COUNT(*)&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  GROUP BY col_colobus_group_size&lt;br /&gt;
  ORDER BY col_colobus_group_size NULLS LAST;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The counts are 18, 3, 11, 1,555, and 49 for codes &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt;, plus 611&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: load the five documented support rows, retain&lt;br /&gt;
the foreign key, make &amp;lt;code&amp;gt;GroupSize&amp;lt;/code&amp;gt; nullable, and preserve source NULL as SQL&lt;br /&gt;
NULL.  The descriptions are &amp;lt;code&amp;gt;group consisted of only 1 col&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;group consisted&lt;br /&gt;
of only 2 col&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;group defined as small by observer&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;group but no size&lt;br /&gt;
estimate given&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;group defined as large&amp;lt;/code&amp;gt;.  This excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#209) COLOBUS focal identifiers require approved corrections ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Five source rows use four noncanonical focal identifiers.  The approved&lt;br /&gt;
corrections are &amp;lt;code&amp;gt;GRP&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;STR&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;FLT&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;FLI&amp;lt;/code&amp;gt;.  &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; does not fit the source column&amp;#039;s &amp;lt;code&amp;gt;VARCHAR(3)&amp;lt;/code&amp;gt; declaration.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT col_fol_b_focal_chimp_id, COUNT(*)&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_fol_b_focal_chimp_id IN (&amp;#039;GRP&amp;#039;, &amp;#039;STR&amp;#039;, &amp;#039;U&amp;#039;, &amp;#039;FLT&amp;#039;)&lt;br /&gt;
  GROUP BY col_fol_b_focal_chimp_id&lt;br /&gt;
  ORDER BY col_fol_b_focal_chimp_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved for this conversion: widen only the clean copy of the focal column to&lt;br /&gt;
&amp;lt;code&amp;gt;TEXT&amp;lt;/code&amp;gt;, then apply the four corrections to the five affected rows.  Preserve&lt;br /&gt;
raw, tidy, and easy source fidelity.  This correction itself excludes no&lt;br /&gt;
rows; the two corrected &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; rows are handled independently by Problem #210.&lt;br /&gt;
&lt;br /&gt;
== * (#210) Corrected COLOBUS unknown focals have no valid watch context ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After Problem #209 and all preceding watch loaders, the corrected &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; rows&lt;br /&gt;
at 1982-10-30 09:10 and 1985-07-21 10:12 have neither a B/Other watch nor a&lt;br /&gt;
covering &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; row.  Inferring a community from unrelated same-day&lt;br /&gt;
observations would create unsupported data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT col_encounter_date, col_fol_b_focal_chimp_id, col_begin_time&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_fol_b_focal_chimp_id = &amp;#039;UNK&amp;#039;&lt;br /&gt;
        AND (col_encounter_date, col_begin_time) IN (&lt;br /&gt;
              (DATE &amp;#039;1982-10-30&amp;#039;, TIME &amp;#039;09:10&amp;#039;),&lt;br /&gt;
              (DATE &amp;#039;1985-07-21&amp;#039;, TIME &amp;#039;10:12&amp;#039;))&lt;br /&gt;
  ORDER BY col_encounter_date, col_begin_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The investigator approved excluding exactly these two rows.  Do not create a&lt;br /&gt;
watch with an inferred community.  They do not overlap Problem #201, raising&lt;br /&gt;
the exact exclusion union from 15 to 17 rows.  Of 2,247 source rows, 2,230 are&lt;br /&gt;
therefore eligible for conversion.&lt;br /&gt;
&lt;br /&gt;
== * (#211) Corrected COLOBUS no-focal row has conflicting community context ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After Problem #209 and all preceding watch loaders, the corrected &amp;lt;code&amp;gt;GRP&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row at 1978-09-10 11:33 has neither a B/Other watch nor a covering&lt;br /&gt;
&amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; row.  Its source comment says &amp;lt;code&amp;gt;ON PATROL, KAHAMA. NO FOCAL&amp;lt;/code&amp;gt;, but&lt;br /&gt;
the unique same-day attendance community is &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt; (Kasekela); &amp;lt;code&amp;gt;HK&amp;lt;/code&amp;gt; is Kahama&lt;br /&gt;
and its support notes say that community ended on 1977-12-31.  Either community&lt;br /&gt;
choice would therefore infer unsupported data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT col_encounter_date, col_fol_b_focal_chimp_id,&lt;br /&gt;
       col_begin_time, col_comments&lt;br /&gt;
  FROM clean.colobus&lt;br /&gt;
  WHERE col_fol_b_focal_chimp_id = &amp;#039;NONE&amp;#039;&lt;br /&gt;
        AND col_encounter_date = DATE &amp;#039;1978-09-10&amp;#039;&lt;br /&gt;
        AND col_begin_time = TIME &amp;#039;11:33&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The investigator approved excluding exactly this row.  Do not create a&lt;br /&gt;
no-focal Other watch with an inferred community.  It does not overlap Problems&lt;br /&gt;
#201 or #210, raising the exact exclusion union from 17 to 18 rows.  Of 2,247&lt;br /&gt;
source rows, 2,229 are therefore eligible for conversion.&lt;br /&gt;
&lt;br /&gt;
== (#212) OTHER_SPECIES requires production species support codes ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;codes.species&amp;lt;/code&amp;gt; is empty, while &amp;lt;code&amp;gt;clean.other_species_lookup&amp;lt;/code&amp;gt; contains 15&lt;br /&gt;
unique local-species names and unique English descriptions.  Every&lt;br /&gt;
&amp;lt;code&amp;gt;SPECIES_PRESENT.Species&amp;lt;/code&amp;gt; value must reference a nonempty uppercase&lt;br /&gt;
&amp;lt;code&amp;gt;codes.species.Species&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT UPPER(BTRIM(osl_local_species_name)) AS proposed_species,&lt;br /&gt;
       osl_english_species_name AS proposed_description&lt;br /&gt;
  FROM clean.other_species_lookup&lt;br /&gt;
  ORDER BY proposed_species;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The query returns 15 rows: &amp;lt;code&amp;gt;CHATU&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FUNGO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;KAKAKUONA&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;KENGE&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;KIMA&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;MBOGO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;NGURUWE&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;NKUNGE&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;NYANI&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;NYOKA&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;PONGO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;TUMBILI&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;UNKNOWN&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;UNRECORDED&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;VYONDI&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: populate &amp;lt;code&amp;gt;codes.species&amp;lt;/code&amp;gt; using&lt;br /&gt;
the uppercase, trimmed local lookup name as &amp;lt;code&amp;gt;Species&amp;lt;/code&amp;gt; and the corresponding&lt;br /&gt;
English lookup name as &amp;lt;code&amp;gt;Description&amp;lt;/code&amp;gt;.  Preserve all 15 descriptions exactly,&lt;br /&gt;
including &amp;lt;code&amp;gt;Mongoose? Civet? TBD&amp;lt;/code&amp;gt;.  Sanity must verify the exact 15 pairs and&lt;br /&gt;
their uniqueness before loading.  This excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#213) OTHER_SPECIES species labels require canonical standardization ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The source stores 30 distinct spellings of its classified local-species field.&lt;br /&gt;
Trimmed, case-insensitive matching resolves 6,360 of 6,366 rows to the 15-row&lt;br /&gt;
lookup.  One additional row uses &amp;lt;code&amp;gt;Viondi&amp;lt;/code&amp;gt;, a spelling variant of lookup value&lt;br /&gt;
&amp;lt;code&amp;gt;Vyondi&amp;lt;/code&amp;gt;.  The production code must use the uppercase canonical lookup value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT quote_literal(os_osl_local_species_name) AS source_value,&lt;br /&gt;
       COUNT(*) AS rows&lt;br /&gt;
  FROM clean.other_species AS other_species&lt;br /&gt;
  WHERE NOT EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
            FROM clean.other_species_lookup AS lookup&lt;br /&gt;
            WHERE UPPER(BTRIM(lookup.osl_local_species_name)) =&lt;br /&gt;
                  UPPER(BTRIM(other_species.os_osl_local_species_name)))&lt;br /&gt;
  GROUP BY os_osl_local_species_name&lt;br /&gt;
  ORDER BY source_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The query returns &amp;lt;code&amp;gt;Sokwe mgeni&amp;lt;/code&amp;gt; (4), &amp;lt;code&amp;gt;UMBA &amp;lt;/code&amp;gt; (1), and &amp;lt;code&amp;gt;Viondi&amp;lt;/code&amp;gt; (1).&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: standardize classified source&lt;br /&gt;
labels by trimmed, case-insensitive lookup and use the lookup&amp;#039;s uppercase&lt;br /&gt;
local name as the production code.  Correct the single &amp;lt;code&amp;gt;Viondi&amp;lt;/code&amp;gt; value to&lt;br /&gt;
canonical &amp;lt;code&amp;gt;VYONDI&amp;lt;/code&amp;gt;.  Do not use &amp;lt;code&amp;gt;OS_local_species_name_written&amp;lt;/code&amp;gt; to override&lt;br /&gt;
the classified field.  This resolves 6,361 rows and excludes none.  The five&lt;br /&gt;
remaining unlisted rows are handled independently by Problem #214.&lt;br /&gt;
&lt;br /&gt;
== * (#214) OTHER_SPECIES contains unlisted species labels ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Four source rows are classified as &amp;lt;code&amp;gt;Sokwe mgeni&amp;lt;/code&amp;gt;, and one is classified as&lt;br /&gt;
&amp;lt;code&amp;gt;UMBA &amp;lt;/code&amp;gt; with comment &amp;lt;code&amp;gt;&amp;#039;umba&amp;#039; written on sheet&amp;lt;/code&amp;gt;.  Neither trimmed label exists&lt;br /&gt;
in &amp;lt;code&amp;gt;clean.other_species_lookup&amp;lt;/code&amp;gt; or has an approved &amp;lt;code&amp;gt;codes.species&amp;lt;/code&amp;gt; code or&lt;br /&gt;
description.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT os_fol_date, os_fol_b_focal_animid, os_time_begin,&lt;br /&gt;
       os_time_end, os_osl_local_species_name,&lt;br /&gt;
       os_local_species_name_written, os_comments&lt;br /&gt;
  FROM clean.other_species AS other_species&lt;br /&gt;
  WHERE UPPER(BTRIM(os_osl_local_species_name)) IN (&amp;#039;SOKWE MGENI&amp;#039;, &amp;#039;UMBA&amp;#039;)&lt;br /&gt;
  ORDER BY os_fol_date, os_fol_b_focal_animid, os_time_begin;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The five rows overlap neither the Problem #215 focal exclusion nor the Problem&lt;br /&gt;
#216 time exclusion.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: exclude exactly these five rows at&lt;br /&gt;
the loader boundary.  Do not map either label to &amp;lt;code&amp;gt;UNKNOWN&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;UNRECORDED&amp;lt;/code&amp;gt; and&lt;br /&gt;
do not create unsupported species codes.  These rows overlap none of Problems&lt;br /&gt;
#215, #216, or #219 and contribute five rows to the final 13-row exclusion&lt;br /&gt;
union.&lt;br /&gt;
&lt;br /&gt;
== * (#215) OTHER_SPECIES focal identifiers require normalization and one exclusion ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Twelve rows have trailing spaces in &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, one row uses&lt;br /&gt;
lowercase &amp;lt;code&amp;gt;pl&amp;lt;/code&amp;gt;, and one row uses &amp;lt;code&amp;gt;IMB&amp;lt;/code&amp;gt;.  The trimmed values resolve uniquely&lt;br /&gt;
to canonical production IDs for the 12 edge-spaced rows, and &amp;lt;code&amp;gt;pl&amp;lt;/code&amp;gt; resolves&lt;br /&gt;
uniquely to &amp;lt;code&amp;gt;PL&amp;lt;/code&amp;gt;.  &amp;lt;code&amp;gt;IMB&amp;lt;/code&amp;gt; is absent from &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; and has no watch.&lt;br /&gt;
Although an &amp;lt;code&amp;gt;IMA&amp;lt;/code&amp;gt; B watch exists on the same date, that does not prove the two&lt;br /&gt;
identifiers name the same individual.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT DISTINCT quote_literal(os_fol_b_focal_animid) AS source_focal,&lt;br /&gt;
       array_agg(DISTINCT biography_data.animid&lt;br /&gt;
                 ORDER BY biography_data.animid)&lt;br /&gt;
         FILTER (WHERE biography_data.animid IS NOT NULL) AS canonical_matches&lt;br /&gt;
  FROM clean.other_species&lt;br /&gt;
    LEFT JOIN sokwedb.biography_data&lt;br /&gt;
      ON LOWER(biography_data.animid) =&lt;br /&gt;
         LOWER(BTRIM(other_species.os_fol_b_focal_animid))&lt;br /&gt;
  WHERE os_fol_b_focal_animid &amp;lt;&amp;gt; BTRIM(os_fol_b_focal_animid)&lt;br /&gt;
        OR NOT EXISTS (&lt;br /&gt;
             SELECT 1&lt;br /&gt;
               FROM sokwedb.biography_data AS exact_biography&lt;br /&gt;
               WHERE exact_biography.animid =&lt;br /&gt;
                     other_species.os_fol_b_focal_animid)&lt;br /&gt;
  GROUP BY os_fol_b_focal_animid&lt;br /&gt;
  ORDER BY source_focal;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: resolve trimmed values to the&lt;br /&gt;
single case-insensitive canonical &amp;lt;code&amp;gt;BIOGRAPHY_DATA.AnimID&amp;lt;/code&amp;gt;, preserving its&lt;br /&gt;
stored spelling.  This removes the 12 trailing spaces and maps &amp;lt;code&amp;gt;pl&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;PL&amp;lt;/code&amp;gt;.&lt;br /&gt;
Exclude exactly the &amp;lt;code&amp;gt;IMB&amp;lt;/code&amp;gt; row dated 2016-08-27 at 10:30; do not infer &amp;lt;code&amp;gt;IMA&amp;lt;/code&amp;gt;.&lt;br /&gt;
Normalized source keys remain 6,366 unique.  Exclude exactly the &amp;lt;code&amp;gt;IMB&amp;lt;/code&amp;gt; row;&lt;br /&gt;
the independently missing B watch for &amp;lt;code&amp;gt;FU&amp;lt;/code&amp;gt; is handled by Problem #219.&lt;br /&gt;
&lt;br /&gt;
== * (#216) OTHER_SPECIES event times cannot satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Six source rows cannot be represented without inventing or rewriting event&lt;br /&gt;
times.  Three have NULL end times, two have Start after Stop, and one&lt;br /&gt;
additional row starts before 04:00.  One reversed row also starts after 20:00.&lt;br /&gt;
There are no second-bearing times and no Stop values after 20:00.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT os_fol_date, os_fol_b_focal_animid,&lt;br /&gt;
       os_time_begin, os_time_end, os_duration,&lt;br /&gt;
       os_osl_local_species_name&lt;br /&gt;
  FROM clean.other_species&lt;br /&gt;
  WHERE os_time_end IS NULL&lt;br /&gt;
        OR os_time_begin &amp;gt; os_time_end&lt;br /&gt;
        OR os_time_begin &amp;lt; TIME &amp;#039;04:00&amp;#039;&lt;br /&gt;
        OR os_time_begin &amp;gt; TIME &amp;#039;20:00&amp;#039;&lt;br /&gt;
        OR os_time_end &amp;gt; TIME &amp;#039;20:00&amp;#039;&lt;br /&gt;
  ORDER BY os_fol_date, os_fol_b_focal_animid, os_time_begin;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: exclude the exact six-row union&lt;br /&gt;
at the loader boundary.  Do not derive a missing endpoint from Duration or use&lt;br /&gt;
Duration to replace a recorded time.  Preserve both endpoints for every other&lt;br /&gt;
row; &amp;lt;code&amp;gt;OS&amp;lt;/code&amp;gt; events are intervals and cannot use &amp;lt;code&amp;gt;sdb_no_time&amp;lt;/code&amp;gt;.  These six rows&lt;br /&gt;
do not overlap Problems #214, #215, or #219.  Together with Problems #214 and&lt;br /&gt;
#215 they bring the exclusion union to 12 rows.&lt;br /&gt;
&lt;br /&gt;
== (#217) OTHER_SPECIES comments map to event notes ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;EVENTS.Notes&amp;lt;/code&amp;gt; is NOT NULL.  Of 6,366 source rows, 6,352 have NULL&lt;br /&gt;
&amp;lt;code&amp;gt;OS_comments&amp;lt;/code&amp;gt; and 14 have non-NULL comments.  The separate&lt;br /&gt;
&amp;lt;code&amp;gt;OS_local_species_name_written&amp;lt;/code&amp;gt; field contains free-form audit evidence but&lt;br /&gt;
has no production destination and sometimes differs from the classified&lt;br /&gt;
local-species field.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*) FILTER (WHERE os_comments IS NULL) AS null_comments,&lt;br /&gt;
       COUNT(*) FILTER (WHERE os_comments IS NOT NULL) AS recorded_comments,&lt;br /&gt;
       COUNT(*) FILTER (&lt;br /&gt;
         WHERE os_comments IS NOT NULL&lt;br /&gt;
               AND os_comments &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
               AND BTRIM(os_comments) = &amp;#039;&amp;#039;) AS whitespace_only_comments&lt;br /&gt;
  FROM clean.other_species;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The query returns 6,352 NULL comments, 14 recorded comments, and zero&lt;br /&gt;
nonempty whitespace-only comments.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: map &amp;lt;code&amp;gt;OS_comments&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;EVENTS.Notes&amp;lt;/code&amp;gt;, preserving non-NULL text exactly and mapping SQL NULL to &amp;lt;code&amp;gt;&amp;#039;&amp;#039;&amp;lt;/code&amp;gt;.&lt;br /&gt;
Retain &amp;lt;code&amp;gt;OS_local_species_name_written&amp;lt;/code&amp;gt; in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; as audit evidence; do not&lt;br /&gt;
append it to event notes or use it to override the classified species field.&lt;br /&gt;
This excludes no rows.  Validation must prove two-way parity under&lt;br /&gt;
&amp;lt;code&amp;gt;COALESCE(os_comments, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#218) SPECIES_PRESENT trigger error paths reference nonexistent fields ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Both error-detail paths in&lt;br /&gt;
&amp;lt;code&amp;gt;db/schemas/lib/triggers/create/species_present.m4&amp;lt;/code&amp;gt; reference&lt;br /&gt;
&amp;lt;code&amp;gt;NEW.researchers&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;NEW.nonresearchers&amp;lt;/code&amp;gt;.  Those are &amp;lt;code&amp;gt;HUMANS&amp;lt;/code&amp;gt; columns;&lt;br /&gt;
&amp;lt;code&amp;gt;SPECIES_PRESENT&amp;lt;/code&amp;gt; has only &amp;lt;code&amp;gt;EID&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;Species&amp;lt;/code&amp;gt;.  An invalid related event or a&lt;br /&gt;
conflicting &amp;lt;code&amp;gt;HUMANS&amp;lt;/code&amp;gt; row can therefore obscure the intended integrity error&lt;br /&gt;
with a record-field error.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
grep -nE &amp;#039;NEW\.(researchers|nonresearchers)&amp;#039; \&lt;br /&gt;
  db/schemas/lib/triggers/create/species_present.m4&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The command reports four references, two in each error-detail path.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Resolved on 2026-09-18.  Both invalid attempted-row details now report&lt;br /&gt;
&amp;lt;code&amp;gt;NEW.species&amp;lt;/code&amp;gt;; the related event/watch and conflicting HUMANS details remain.&lt;br /&gt;
The generated trigger was installed in a disposable PostgreSQL 18.6 database,&lt;br /&gt;
and &amp;lt;code&amp;gt;db/tests/species_present_trigger.sql&amp;lt;/code&amp;gt; passed rollback-only tests for wrong&lt;br /&gt;
event behavior, HUMANS conflict, a valid &amp;lt;code&amp;gt;OS&amp;lt;/code&amp;gt; insert, immediate deferred&lt;br /&gt;
constraints, and fixture cleanup.  This is a trigger repair and excludes no&lt;br /&gt;
source rows.&lt;br /&gt;
&lt;br /&gt;
== * (#219) OTHER_SPECIES row has no B watch at the required load stage ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After the normal conversion sequence through &amp;lt;code&amp;gt;load_follow_to_watches&amp;lt;/code&amp;gt;, the&lt;br /&gt;
otherwise valid &amp;lt;code&amp;gt;FU&amp;lt;/code&amp;gt; row dated 2015-12-13 at 16:45 has no B watch.  &amp;lt;code&amp;gt;FU&amp;lt;/code&amp;gt; has no&lt;br /&gt;
source FOLLOW row that day.  Later aggression and grooming loaders create a B&lt;br /&gt;
watch for the same focal/date, but moving OTHER_SPECIES later would make its&lt;br /&gt;
prerequisite depend on unrelated observations and conceal the stage-specific&lt;br /&gt;
source gap.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT os_fol_date, os_fol_b_focal_animid, os_time_begin, os_time_end,&lt;br /&gt;
       os_osl_local_species_name, os_comments&lt;br /&gt;
  FROM clean.other_species&lt;br /&gt;
  WHERE os_fol_date = DATE &amp;#039;2015-12-13&amp;#039;&lt;br /&gt;
        AND os_fol_b_focal_animid = &amp;#039;FU&amp;#039;&lt;br /&gt;
        AND os_time_begin = TIME &amp;#039;16:45&amp;#039;;&lt;br /&gt;
&lt;br /&gt;
SELECT watches.wid&lt;br /&gt;
  FROM sokwedb.watches&lt;br /&gt;
  WHERE watches.animid = &amp;#039;FU&amp;#039;&lt;br /&gt;
        AND watches.date = DATE &amp;#039;2015-12-13&amp;#039;&lt;br /&gt;
        AND watches.type = &amp;#039;B&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At the required post-follow loader stage, the first query returns the one&lt;br /&gt;
&amp;lt;code&amp;gt;VYONDI&amp;lt;/code&amp;gt; interval from 16:45 through 17:00 and the second query returns no row.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: exclude exactly the &amp;lt;code&amp;gt;FU&amp;lt;/code&amp;gt; row dated&lt;br /&gt;
2015-12-13 at 16:45.  Keep OTHER_SPECIES immediately after&lt;br /&gt;
&amp;lt;code&amp;gt;load_follow_to_watches&amp;lt;/code&amp;gt;; do not create a watch and do not reuse a watch that&lt;br /&gt;
only appears after later aggression or grooming loads.  This row overlaps&lt;br /&gt;
none of Problems #214, #215, or #216, increasing the final exclusion union to&lt;br /&gt;
13 rows and leaving 6,353 rows eligible for conversion.&lt;br /&gt;
&lt;br /&gt;
== (#220) FOLLOW_MAP_LOCATION rows can contain two location representations ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Of 1,823,766 &amp;lt;code&amp;gt;FOLLOW_MAP_LOCATION&amp;lt;/code&amp;gt; rows, 703,807 contain both a paper map&lt;br /&gt;
sequence and paired UTM coordinates.  Production stores PAPER and UTM details&lt;br /&gt;
on separate EVENTS rows, so treating the source row as a choice of one detail&lt;br /&gt;
table would discard one recorded representation.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
The read-only &amp;lt;code&amp;gt;conversion/follow_map_location_profile.sql&amp;lt;/code&amp;gt; reports 703,807&lt;br /&gt;
dual rows, 69 paper-only rows, 1,119,890 UTM-only rows, and no row with neither&lt;br /&gt;
representation.  Before exclusions this is 703,876 PAPER representations and&lt;br /&gt;
1,823,697 UTM representations.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: emit one PAPER event and one UTM&lt;br /&gt;
event for every eligible dual row.  Both events use the same location watch&lt;br /&gt;
and source time.  No loader may choose one representation or infer that UTM&lt;br /&gt;
coordinates accompanying a paper map are direct GPS observations.&lt;br /&gt;
&lt;br /&gt;
== * (#221) FOLLOW_MAP_LOCATION focal identities are missing or unsupported ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Location watches require a real &amp;lt;code&amp;gt;BIOGRAPHY_DATA.AnimID&amp;lt;/code&amp;gt;; &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; is prohibited.&lt;br /&gt;
The refreshed source has 6,509 NULL-focal rows and 7,495 rows in 33 trimmed&lt;br /&gt;
labels having no normalized biography match.  The source does not contain&lt;br /&gt;
evidence from which to invent their identities.&lt;br /&gt;
&lt;br /&gt;
Fourteen noncanonical spellings covering 609 rows do resolve uniquely under&lt;br /&gt;
&amp;lt;code&amp;gt;lower(normalize(btrim(value)))&amp;lt;/code&amp;gt;.  They include edge-space variants and the&lt;br /&gt;
case variants &amp;lt;code&amp;gt;Nas&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;NAS&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fd&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FD&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;pl&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;PL&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;sw &amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;SW&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Run &amp;lt;code&amp;gt;conversion/follow_map_location_profile.sql&amp;lt;/code&amp;gt;.  It prints every unmatched&lt;br /&gt;
trimmed label with its row count and date range, and reports 214 raw non-NULL&lt;br /&gt;
spellings, 204 trimmed spellings, and 200 normalized spelling classes.  The&lt;br /&gt;
unmatched labels include punctuation-prefixed labels, apparent suffixed IDs,&lt;br /&gt;
and labels that cannot safely be corrected from spelling alone.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: standardize the 609 uniquely&lt;br /&gt;
matched rows to the stored biography spelling and exclude all 6,509 NULL-focal&lt;br /&gt;
rows and 7,495 unmatched-label rows.  This excludes 14,004 rows before overlap&lt;br /&gt;
with other problems.  Do not derive a focal from coordinates, community, date,&lt;br /&gt;
another observation, or a nearby follow.&lt;br /&gt;
&lt;br /&gt;
== * (#222) FOLLOW_MAP_LOCATION event times cannot be represented ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PAPER and UTM events require a recorded minute from 04:00 through 20:00.  The&lt;br /&gt;
refreshed source has 25,921 NULL times, 676 times before 04:00, and 258 after&lt;br /&gt;
20:00.  One recorded time is midnight, so midnight cannot represent NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/follow_map_location_profile.sql&amp;lt;/code&amp;gt; reports the exact classes and&lt;br /&gt;
representation impact.  NULL times affect 25,921 PAPER and 25,914 UTM&lt;br /&gt;
representations; early times affect 676 UTM representations; late times&lt;br /&gt;
affect 7 PAPER and 258 UTM representations.  Recorded seconds are always zero.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: exclude all 25,921 NULL-time rows,&lt;br /&gt;
676 rows before 04:00, and 258 rows after 20:00.  This excludes 26,855 rows&lt;br /&gt;
before overlap with other problems.  Do not map NULL to midnight, clamp&lt;br /&gt;
out-of-range times, round times, or derive a time from neighboring source rows.&lt;br /&gt;
&lt;br /&gt;
== * (#223) FOLLOW_MAP_LOCATION event keys collide ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production allows at most one PAPER and one UTM event for a location watch at&lt;br /&gt;
a given time.  Before focal normalization and after omitting NULL times, the&lt;br /&gt;
source has 23,801 colliding PAPER focal/date/time keys and 25,561 colliding UTM&lt;br /&gt;
keys.  Generated IDs cannot make those keys representable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
The profile reports 47,603 PAPER rows in colliding keys, with 23,802 excess&lt;br /&gt;
rows, and 51,255 UTM rows, with 25,694 excess rows.  Only 75 PAPER keys (150&lt;br /&gt;
rows) and 150 UTM keys (300 rows) have one identical representation value;&lt;br /&gt;
23,726 PAPER keys and 25,411 UTM keys contain genuinely different recorded&lt;br /&gt;
representations.  One UTM key has 14 source rows.&lt;br /&gt;
&lt;br /&gt;
The complete source additionally has 151 exact duplicate pairs across all 14&lt;br /&gt;
columns.  A generated conversion ID does not establish that either copy is a&lt;br /&gt;
separate production observation.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: exclude every source row whose&lt;br /&gt;
canonical focal/date/time key collides for either available representation.&lt;br /&gt;
This excludes the entire source row, rather than retaining half of a dual row,&lt;br /&gt;
and removes 51,457 rows before overlap with other problems.  Do not add&lt;br /&gt;
arbitrary time offsets, select a row by generated order, or let &amp;lt;code&amp;gt;ON CONFLICT&amp;lt;/code&amp;gt;&lt;br /&gt;
discard rows.&lt;br /&gt;
&lt;br /&gt;
== * (#224) FOLLOW_MAP_LOCATION community recovery counts depend on focal normalization ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The investigator approved dated &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; as canonical, &amp;lt;code&amp;gt;Unknown&amp;lt;/code&amp;gt; only for&lt;br /&gt;
a NULL source community without dated membership, and exclusion of unsupported&lt;br /&gt;
non-NULL communities and valid non-NULL communities without matching dated&lt;br /&gt;
membership or disagreeing with membership.&lt;br /&gt;
&lt;br /&gt;
The previously approved exact counts used a case-sensitive&lt;br /&gt;
&amp;lt;code&amp;gt;btrim(focal) = COMM_MEMBS.AnimID&amp;lt;/code&amp;gt; join.  The loader specification instead&lt;br /&gt;
requires canonical focal matching by &amp;lt;code&amp;gt;lower(normalize(btrim(value)))&amp;lt;/code&amp;gt;.&lt;br /&gt;
Applying that rule resolves 118 additional rows through &amp;lt;code&amp;gt;Nas&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;NAS&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fd&amp;lt;/code&amp;gt;&lt;br /&gt;
to &amp;lt;code&amp;gt;FD&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;pl&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;PL&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;sw &amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;SW&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
On the refreshed source, canonical matching reports:&lt;br /&gt;
&lt;br /&gt;
* 1,095,525 NULL-source rows with membership;&lt;br /&gt;
* 14,412 NULL-source rows without membership;&lt;br /&gt;
* 712,457 valid-source rows agreeing with membership;&lt;br /&gt;
* 245 valid-source rows without membership;&lt;br /&gt;
* 1,105 valid-source rows disagreeing with membership; and&lt;br /&gt;
* 22 unsupported non-NULL source communities.&lt;br /&gt;
&lt;br /&gt;
The canonical result would exclude 1,372 rows rather than 1,455.  The old&lt;br /&gt;
trim-only result remains exactly 1,095,490, 14,447, 712,374, 328, 1,105, and&lt;br /&gt;
22 respectively.  The 83-row exclusion difference is 51 &amp;lt;code&amp;gt;fd&amp;lt;/code&amp;gt;, 14 &amp;lt;code&amp;gt;pl&amp;lt;/code&amp;gt;, and 18&lt;br /&gt;
&amp;lt;code&amp;gt;sw &amp;lt;/code&amp;gt; rows whose valid source community agrees with canonical membership.  The&lt;br /&gt;
other 35 newly resolved rows are NULL-community &amp;lt;code&amp;gt;Nas&amp;lt;/code&amp;gt; rows and would change&lt;br /&gt;
from &amp;lt;code&amp;gt;Unknown&amp;lt;/code&amp;gt; to membership community &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Canonical focal standardization was approved under Problem #221.  The final&lt;br /&gt;
Problem #224 exclusion union is therefore 1,372 rows: 22 unsupported non-NULL&lt;br /&gt;
communities, 245 valid communities without dated membership, and 1,105 valid&lt;br /&gt;
communities disagreeing with membership.  Do not preserve the old counts by&lt;br /&gt;
using a less canonical identity join.&lt;br /&gt;
&lt;br /&gt;
== (#225) LOCATION_ORIGINS lacks approved support values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;LOCATION_ORIGINS&amp;lt;/code&amp;gt; is empty, but every PAPER and UTM detail requires an origin.&lt;br /&gt;
The source contains &amp;lt;code&amp;gt;GPS&amp;lt;/code&amp;gt; on 1,113,996 rows, &amp;lt;code&amp;gt;Paper Map&amp;lt;/code&amp;gt; on 15,165 rows,&lt;br /&gt;
&amp;lt;code&amp;gt;paper maps&amp;lt;/code&amp;gt; on 27,273 rows, and NULL on 667,332 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
The profile reports that NULL origin affects 661,438 paper representations and&lt;br /&gt;
667,332 UTM representations.  Of the latter, 5,894 are UTM-only rows.  A UTM&lt;br /&gt;
representation therefore does not prove GPS origin, and dual historical rows&lt;br /&gt;
may contain coordinates converted from a paper map.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18.  Populate canonical origins &amp;lt;code&amp;gt;GPS&amp;lt;/code&amp;gt;&lt;br /&gt;
(&amp;lt;code&amp;gt;GPS&amp;lt;/code&amp;gt;), &amp;lt;code&amp;gt;PAPER&amp;lt;/code&amp;gt; (&amp;lt;code&amp;gt;Paper map&amp;lt;/code&amp;gt;), and &amp;lt;code&amp;gt;UNKNOWN&amp;lt;/code&amp;gt; (&amp;lt;code&amp;gt;Unknown&amp;lt;/code&amp;gt;).  The uppercase code&lt;br /&gt;
is required by the production support-table constraint.  Map source &amp;lt;code&amp;gt;GPS&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;GPS&amp;lt;/code&amp;gt;, both &amp;lt;code&amp;gt;Paper Map&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;paper maps&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;PAPER&amp;lt;/code&amp;gt;, and SQL NULL to &amp;lt;code&amp;gt;UNKNOWN&amp;lt;/code&amp;gt;.&lt;br /&gt;
Apply the same source mapping to both representations; do not infer GPS solely&lt;br /&gt;
from the presence of coordinates.&lt;br /&gt;
&lt;br /&gt;
== (#226) FOLLOW_MAP_LOCATION entered-person support is incomplete ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;Entered&amp;lt;/code&amp;gt; is NULL on 1,781,328 rows.  Recorded values include two labels for&lt;br /&gt;
Danielle Lodge and the full name &amp;lt;code&amp;gt;Julianna Turner&amp;lt;/code&amp;gt;; those values did not&lt;br /&gt;
previously have complete PEOPLE support.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
The refreshed profile reports 17,794 &amp;lt;code&amp;gt;DL&amp;lt;/code&amp;gt;, 8,902 &amp;lt;code&amp;gt;Danielle Lodge&amp;lt;/code&amp;gt;, 7,968&lt;br /&gt;
&amp;lt;code&amp;gt;Julianna Turner&amp;lt;/code&amp;gt;, 2,314 &amp;lt;code&amp;gt;KS&amp;lt;/code&amp;gt;, 4,806 &amp;lt;code&amp;gt;SM&amp;lt;/code&amp;gt;, and 654 &amp;lt;code&amp;gt;SR&amp;lt;/code&amp;gt; rows.  The refreshed&lt;br /&gt;
support table contains &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;KS&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;SM&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;SR&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;DL&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;Julianna Turner&amp;lt;/code&amp;gt;&lt;br /&gt;
with no normalized duplicate.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18.  Map NULL to &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;; map both &amp;lt;code&amp;gt;DL&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;Danielle Lodge&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;DL&amp;lt;/code&amp;gt;; map &amp;lt;code&amp;gt;Julianna Turner&amp;lt;/code&amp;gt; to that same-named code;&lt;br /&gt;
and preserve &amp;lt;code&amp;gt;KS&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;SM&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;SR&amp;lt;/code&amp;gt;.  Match with&lt;br /&gt;
&amp;lt;code&amp;gt;lower(normalize(btrim(value)))&amp;lt;/code&amp;gt;.  The manually maintained &amp;lt;code&amp;gt;clean.people&amp;lt;/code&amp;gt; seed&lt;br /&gt;
contains &amp;lt;code&amp;gt;DL&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;Julianna Turner&amp;lt;/code&amp;gt;, both described as follow map location data&lt;br /&gt;
entry.  This decision excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#227) FOLLOW_MAP_LOCATION contains spatial anomalies ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Thirty-nine UTM rows have coordinates outside the expected local orientation,&lt;br /&gt;
and two rows have elevation 20,000.  The production numeric limits equal the&lt;br /&gt;
source extrema, so constraint acceptance is not evidence of geographic&lt;br /&gt;
validity.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/follow_map_location_profile.sql&amp;lt;/code&amp;gt; prints all 41 anomaly rows.  The&lt;br /&gt;
39 coordinate rows include equal axes, axes both near 9.48 million, and values&lt;br /&gt;
that appear to have missing digits.  The two elevation rows are dated&lt;br /&gt;
2020-03-08 (&amp;lt;code&amp;gt;KOM&amp;lt;/code&amp;gt;) and 2021-06-06 (&amp;lt;code&amp;gt;AP&amp;lt;/code&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: preserve all 39 unusual coordinate&lt;br /&gt;
rows and both elevation-20,000 rows exactly as recorded.  They satisfy current&lt;br /&gt;
production bounds.  Do not swap axes, add digits, replace elevation by&lt;br /&gt;
heuristic, or exclude these rows solely because they are anomalous.&lt;br /&gt;
&lt;br /&gt;
== (#228) FOLLOW_MAP_LOCATION has NULL optional metadata ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;FollowNum&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;Notes&amp;lt;/code&amp;gt; are NOT NULL but allow empty text.  Source&lt;br /&gt;
&amp;lt;code&amp;gt;FollowNum&amp;lt;/code&amp;gt; is NULL on 1,163,270 rows and &amp;lt;code&amp;gt;Notes&amp;lt;/code&amp;gt; is NULL on 1,823,410 rows.&lt;br /&gt;
&amp;lt;code&amp;gt;fml_update&amp;lt;/code&amp;gt; is populated on 48,796 rows but has no production destination.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
The profile reports 18,541 distinct populated follow numbers of length 5&lt;br /&gt;
through 10, and 356 populated notes containing 197 distinct values.  Every&lt;br /&gt;
populated value can be preserved exactly.  Every populated &amp;lt;code&amp;gt;fml_update&amp;lt;/code&amp;gt; is&lt;br /&gt;
2018-05-20.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: map SQL NULL &amp;lt;code&amp;gt;FollowNum&amp;lt;/code&amp;gt; and&lt;br /&gt;
&amp;lt;code&amp;gt;Notes&amp;lt;/code&amp;gt; to empty text while preserving populated values exactly.  Retain&lt;br /&gt;
&amp;lt;code&amp;gt;fml_update&amp;lt;/code&amp;gt; only in the clean audit source and do not append it to notes.&lt;br /&gt;
&lt;br /&gt;
== * (#229) FOLLOW_MAP_LOCATION dates exceed focal departure dates ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;WATCHES&amp;lt;/code&amp;gt; rejects a watch before the focal&amp;#039;s birth or entry and after the&lt;br /&gt;
focal&amp;#039;s departure from study.  The source has 229 canonically matched rows&lt;br /&gt;
after &amp;lt;code&amp;gt;BIOGRAPHY_DATA.DepartDate&amp;lt;/code&amp;gt;, of which 47 are already excluded by both&lt;br /&gt;
Problems #222 and #224.  The approved Problems #220-#228 projection therefore&lt;br /&gt;
still contained 182 rows after departure.  There are no projected rows before&lt;br /&gt;
birth or entry.&lt;br /&gt;
&lt;br /&gt;
All 182 affected rows are UTM-only GPS records with NULL source community,&lt;br /&gt;
NULL entered person, and valid event times.  They form two location watches:&lt;br /&gt;
&lt;br /&gt;
* focal &amp;lt;code&amp;gt;S&amp;lt;/code&amp;gt;, 134 rows on 2017-11-29 from 06:46 through 18:42, after departure&lt;br /&gt;
  date 1968-01-19; and&lt;br /&gt;
* focal &amp;lt;code&amp;gt;DE&amp;lt;/code&amp;gt;, 48 rows on 2017-12-27 from 12:41 through 17:02, after departure&lt;br /&gt;
  date 1974-05-07.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT source.*,&lt;br /&gt;
       biography_data.birthdate,&lt;br /&gt;
       biography_data.entrydate,&lt;br /&gt;
       biography_data.departdate&lt;br /&gt;
  FROM clean.follow_map_location AS source&lt;br /&gt;
  JOIN sokwedb.biography_data&lt;br /&gt;
    ON lower(normalize(biography_data.animid)) =&lt;br /&gt;
         lower(normalize(btrim(source.fml_fol_b_focal_animid)))&lt;br /&gt;
 WHERE source.fml_fol_date &amp;lt; biography_data.birthdate&lt;br /&gt;
       OR source.fml_fol_date &amp;lt; biography_data.entrydate&lt;br /&gt;
       OR source.fml_fol_date &amp;gt; biography_data.departdate&lt;br /&gt;
 ORDER BY source.fml_conversion_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The first full chunked load stopped when inserting the &amp;lt;code&amp;gt;S&amp;lt;/code&amp;gt; location watch on&lt;br /&gt;
2017-11-29 because the &amp;lt;code&amp;gt;WATCHES&amp;lt;/code&amp;gt; trigger reported its 1968-01-19 departure.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: exclude all post-departure rows.&lt;br /&gt;
The predicate identifies 229 raw rows; 47 overlap both Problems #222 and #224,&lt;br /&gt;
so Problem #229 adds exactly 182 rows to the exclusion union.  The final&lt;br /&gt;
distinct union is 93,470, leaving 1,730,296 eligible source rows.  Do not weaken&lt;br /&gt;
the production biography-date trigger, substitute another focal, or alter the&lt;br /&gt;
recorded location dates.&lt;br /&gt;
&lt;br /&gt;
== (#230) Location events lack a matching focal arrival interval ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production documentation expects location times to fall within a focal&amp;#039;s&lt;br /&gt;
arrival interval on a B watch, but the database reports this condition as a&lt;br /&gt;
warning rather than rejecting the event.  Of 2,358,936 loaded PAPER and UTM&lt;br /&gt;
events, 779,682 do not fall within a matching focal arrival interval.&lt;br /&gt;
&lt;br /&gt;
The refreshed post-load profile separates them as follows:&lt;br /&gt;
&lt;br /&gt;
* 724,075 events have no B watch for the focal/date: 196 PAPER and 723,879 UTM;&lt;br /&gt;
* 404 events have a B watch but no ARR interval: 202 PAPER and 202 UTM; and&lt;br /&gt;
* 55,203 events are outside every ARR interval on an existing B watch: 19,118&lt;br /&gt;
  PAPER and 36,085 UTM.&lt;br /&gt;
&lt;br /&gt;
These classes cover 1973-01-04 through 2022-11-28.  The large no-B-watch UTM&lt;br /&gt;
class reflects location coverage continuing well beyond the consolidated&lt;br /&gt;
follow data; it is not evidence from which to invent a follow or arrival time.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Join location &amp;lt;code&amp;gt;EVENTS&amp;lt;/code&amp;gt; and their &amp;lt;code&amp;gt;L&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;WATCHES&amp;lt;/code&amp;gt; rows to B watches on focal/date,&lt;br /&gt;
then test whether the location start falls between an ARR event&amp;#039;s start and&lt;br /&gt;
stop.  &amp;lt;code&amp;gt;conversion/load_follow_map_location_validation.sql&amp;lt;/code&amp;gt; reports the exact&lt;br /&gt;
779,682-event total after proving source-to-target parity.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-18: preserve all 779,682 events as&lt;br /&gt;
recorded.  The production schema permits them and the arrival relationship is&lt;br /&gt;
documentation guidance enforced as a warning.  Validation asserts the exact&lt;br /&gt;
count and continues to report it.  Do not create B watches, manufacture ARR&lt;br /&gt;
intervals, alter location times, or exclude otherwise eligible source rows to&lt;br /&gt;
satisfy this guidance.&lt;br /&gt;
&lt;br /&gt;
== * (#231) CHIMPS_ON_TIKIS source identifiers lack biography support ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Three &amp;lt;code&amp;gt;clean.chimps_on_tikis_lookup&amp;lt;/code&amp;gt; rows cannot satisfy the required&lt;br /&gt;
&amp;lt;code&amp;gt;CHIMPS_ON_TIKIS.AnimID&amp;lt;/code&amp;gt; foreign key: &amp;lt;code&amp;gt;MGENI&amp;lt;/code&amp;gt; occurs once in &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt; and once in&lt;br /&gt;
&amp;lt;code&amp;gt;MT&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;VAN mtoto&amp;lt;/code&amp;gt; occurs once in &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt;.  &amp;lt;code&amp;gt;MGENI&amp;lt;/code&amp;gt; does not identify a&lt;br /&gt;
particular production individual.  &amp;lt;code&amp;gt;VAN mtoto&amp;lt;/code&amp;gt; is not a production AnimID.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;VIS&amp;lt;/code&amp;gt; is VAN&amp;#039;s only supported child alive on the Tiki creation and deployment&lt;br /&gt;
dates and had &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt; membership, but that evidence does not establish that the&lt;br /&gt;
source label denotes &amp;lt;code&amp;gt;VIS&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT source.*&lt;br /&gt;
  FROM clean.chimps_on_tikis_lookup AS source&lt;br /&gt;
  LEFT JOIN sokwedb.biography_data AS biography&lt;br /&gt;
    ON biography.animid = source.b_animid&lt;br /&gt;
  WHERE biography.animid IS NULL&lt;br /&gt;
  ORDER BY source.tiki_creation_date,&lt;br /&gt;
           source.community_id,&lt;br /&gt;
           source.b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns exactly:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
 b_animid  | tiki_creation_date | community_id | tiki_deployment_date | new_addition&lt;br /&gt;
-----------+--------------------+--------------+----------------------+-------------&lt;br /&gt;
 MGENI     | 2020-08-04         | KK           | 2020-08-20           | N&lt;br /&gt;
 VAN mtoto | 2020-08-04         | KK           | 2020-08-20           | Y&lt;br /&gt;
 MGENI     | 2020-08-04         | MT           | 2020-08-20           | N&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: exclude exactly these three&lt;br /&gt;
complete source tuples.  Do not create a biography row for &amp;lt;code&amp;gt;MGENI&amp;lt;/code&amp;gt;, infer an&lt;br /&gt;
individual for either &amp;lt;code&amp;gt;MGENI&amp;lt;/code&amp;gt; row, or map &amp;lt;code&amp;gt;VAN mtoto&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;VIS&amp;lt;/code&amp;gt;.  The loader&lt;br /&gt;
sanity check must prove the exact exclusion inventory and reject any remaining&lt;br /&gt;
unsupported identifier.  The dated source has 84 rows, leaving 81 eligible&lt;br /&gt;
rows.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  The shared projection compares all&lt;br /&gt;
five source columns with &amp;lt;code&amp;gt;IS NOT DISTINCT FROM&amp;lt;/code&amp;gt; and excludes only the approved&lt;br /&gt;
tuples.  Read-only sanity reproduced 84 source rows, the exact three-row&lt;br /&gt;
exclusion multiset, and 81 supported unique projected rows.  The direct load&lt;br /&gt;
produced 81 target rows, including 13 &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt; and 68 &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt; newcomer values;&lt;br /&gt;
bidirectional &amp;lt;code&amp;gt;EXCEPT ALL&amp;lt;/code&amp;gt; returned no payload differences.  An intentional&lt;br /&gt;
post-insert failure rolled all 81 rows back to zero.  An approved clean retry&lt;br /&gt;
reproduced the same payload digest, &amp;lt;code&amp;gt;28e17175eba7622cbf9afc7130dc1900&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#232) CHIMPS_ON_TIKIS deployment-date index targets AnimID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The index named &amp;lt;code&amp;gt;chimps_on_tikis_tikideployementdate&amp;lt;/code&amp;gt; is defined on &amp;lt;code&amp;gt;animid&amp;lt;/code&amp;gt;&lt;br /&gt;
instead of &amp;lt;code&amp;gt;tikideploymentdate&amp;lt;/code&amp;gt;.  It duplicates the separate AnimID index and&lt;br /&gt;
leaves the documented deployment-date lookup unindexed.  The index name also&lt;br /&gt;
misspells &amp;quot;deployment&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT indexname, indexdef&lt;br /&gt;
  FROM pg_indexes&lt;br /&gt;
  WHERE schemaname = &amp;#039;housekeeping&amp;#039;&lt;br /&gt;
        AND tablename = &amp;#039;chimps_on_tikis&amp;#039;&lt;br /&gt;
        AND indexname = &amp;#039;chimps_on_tikis_tikideployementdate&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The live PostgreSQL 18.6 database reports an index using &amp;lt;code&amp;gt;btree (animid)&amp;lt;/code&amp;gt;.&lt;br /&gt;
The owning definition is in&lt;br /&gt;
&amp;lt;code&amp;gt;db/schemas/housekeeping/indexes/create/chimps_on_tikis.m4&amp;lt;/code&amp;gt;; the corresponding&lt;br /&gt;
drop name is in the neighboring drop M4.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: before loading this table, replace&lt;br /&gt;
the defective index with correctly spelled&lt;br /&gt;
&amp;lt;code&amp;gt;chimps_on_tikis_tikideploymentdate&amp;lt;/code&amp;gt; on &amp;lt;code&amp;gt;(tikideploymentdate)&amp;lt;/code&amp;gt;.  Update both&lt;br /&gt;
owning M4 files, regenerate owned index SQL artifacts, and verify the live&lt;br /&gt;
index definition in a disposable database.  This schema prerequisite excludes&lt;br /&gt;
no source rows.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  The owning create and drop M4 files&lt;br /&gt;
now use the corrected index, while the drop definition also removes the legacy&lt;br /&gt;
misspelled name during an in-place migration.  Repository generation rebuilt&lt;br /&gt;
the per-object and aggregate index SQL.  Both the focused local migration and&lt;br /&gt;
the freshly rebuilt disposable integration schema reported exactly one index&lt;br /&gt;
on &amp;lt;code&amp;gt;tikideploymentdate&amp;lt;/code&amp;gt;, named&lt;br /&gt;
&amp;lt;code&amp;gt;chimps_on_tikis_tikideploymentdate&amp;lt;/code&amp;gt;, and no legacy index.&lt;br /&gt;
&lt;br /&gt;
== * (#233) SIV_STATUS_BOUT comma-bearing names shifted into later columns ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Five source rows have a comma-bearing original name split across&lt;br /&gt;
&amp;lt;code&amp;gt;SSB_names_used&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;SSB_in_bio_table&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;SSB_notes&amp;lt;/code&amp;gt;.  The leading name&lt;br /&gt;
fragment is in &amp;lt;code&amp;gt;SSB_names_used&amp;lt;/code&amp;gt;, the trailing fragment is in&lt;br /&gt;
&amp;lt;code&amp;gt;SSB_in_bio_table&amp;lt;/code&amp;gt;, and the actual biography flag &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; is in &amp;lt;code&amp;gt;SSB_notes&amp;lt;/code&amp;gt;.&lt;br /&gt;
These malformed values already occur in &amp;lt;code&amp;gt;easy.siv_status_bout&amp;lt;/code&amp;gt;; they were not&lt;br /&gt;
introduced by a clean-stage transformation.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT source.*&lt;br /&gt;
  FROM clean.siv_status_bout AS source&lt;br /&gt;
  WHERE source.ssb_in_bio_table NOT IN (&amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;)&lt;br /&gt;
  ORDER BY source.ssb_b_animid, source.ssb_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns exactly five rows: &amp;lt;code&amp;gt;DB&amp;lt;/code&amp;gt; with fragments &amp;lt;code&amp;gt;&amp;quot;Darbee&amp;lt;/code&amp;gt; and&lt;br /&gt;
&amp;lt;code&amp;gt;NA&amp;quot;&amp;lt;/code&amp;gt;; &amp;lt;code&amp;gt;TOF&amp;lt;/code&amp;gt; with &amp;lt;code&amp;gt;&amp;quot;Tofiki&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;gone&amp;quot;&amp;lt;/code&amp;gt;; and all three &amp;lt;code&amp;gt;TT&amp;lt;/code&amp;gt; bouts with&lt;br /&gt;
&amp;lt;code&amp;gt;&amp;quot;Kati (Tita)&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;Kati-tita&amp;quot;&amp;lt;/code&amp;gt;.  Each row has &amp;lt;code&amp;gt;SSB_notes = &amp;#039;1&amp;#039;&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: exclude these five exact complete&lt;br /&gt;
source tuples.  The final decision supersedes an earlier consideration of&lt;br /&gt;
reconstructing the values in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt;.  Do not repair the imported fields,&lt;br /&gt;
infer punctuation, or use a broad identifier-only exclusion.  Sanity must&lt;br /&gt;
prove the exact five-row multiset, and the canonical projection must compare&lt;br /&gt;
all ten source columns with NULL-safe equality.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  &amp;lt;code&amp;gt;conversion/siv_status_bout_projection.sql&amp;lt;/code&amp;gt;&lt;br /&gt;
encodes the exact five-row &amp;lt;code&amp;gt;DB&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;TOF&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;TT&amp;lt;/code&amp;gt; tuple set, including the two&lt;br /&gt;
literal embedded double-quote characters recovered with &amp;lt;code&amp;gt;quote_literal&amp;lt;/code&amp;gt;.&lt;br /&gt;
Read-only sanity, run against the local PostgreSQL 18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; database,&lt;br /&gt;
reproduced this exact multiset with bidirectional &amp;lt;code&amp;gt;EXCEPT ALL&amp;lt;/code&amp;gt; and confirmed&lt;br /&gt;
it overlaps Problem #234 in exactly one row and Problem #235 in none.&lt;br /&gt;
&lt;br /&gt;
== * (#234) SIV_STATUS_BOUT required test dates are missing ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Twelve source rows cannot satisfy the target&amp;#039;s non-NULL &amp;lt;code&amp;gt;FirstTestDate&amp;lt;/code&amp;gt; and&lt;br /&gt;
&amp;lt;code&amp;gt;LastTestDate&amp;lt;/code&amp;gt; columns.  All 11 &amp;lt;code&amp;gt;unk&amp;lt;/code&amp;gt; bouts have both dates absent.  The &amp;lt;code&amp;gt;VID&amp;lt;/code&amp;gt;&lt;br /&gt;
positive bout has &amp;lt;code&amp;gt;FirstTestDate = 2000-09-09&amp;lt;/code&amp;gt; and a NULL &amp;lt;code&amp;gt;LastTestDate&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT source.*&lt;br /&gt;
  FROM clean.siv_status_bout AS source&lt;br /&gt;
  WHERE source.ssb_first_test_date IS NULL&lt;br /&gt;
     OR source.ssb_last_test_date IS NULL&lt;br /&gt;
  ORDER BY source.ssb_b_animid, source.ssb_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 11 unknown bouts belong to &amp;lt;code&amp;gt;BIM&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;CH100&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;ERI&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GIM&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;HAI&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;MAK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;RUD&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;TT&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;YD&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;ZS&amp;lt;/code&amp;gt;.  The &amp;lt;code&amp;gt;TT&amp;lt;/code&amp;gt; unknown bout is also one of the&lt;br /&gt;
five Problem #233 shifted rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: exclude the 12 exact complete&lt;br /&gt;
source tuples.  Do not fill absent tests from bout boundaries and do not change&lt;br /&gt;
target nullability.  Match all ten source columns with NULL-safe equality.&lt;br /&gt;
Because the &amp;lt;code&amp;gt;TT&amp;lt;/code&amp;gt; row overlaps Problem #233, count it once in the combined&lt;br /&gt;
exclusion set.&lt;br /&gt;
&lt;br /&gt;
This decision removes every source &amp;lt;code&amp;gt;unk&amp;lt;/code&amp;gt; bout.  Ten retained positive rows&lt;br /&gt;
therefore keep &amp;lt;code&amp;gt;BoutNumber = 3&amp;lt;/code&amp;gt; without a loaded bout 2.  Preserve those source&lt;br /&gt;
bout numbers; do not renumber or synthesize replacement bouts.  Update the&lt;br /&gt;
production documentation to report the resulting snapshot accurately.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  The projection excludes the exact&lt;br /&gt;
12-row tuple set, including the &amp;lt;code&amp;gt;TT&amp;lt;/code&amp;gt; row shared with Problem #233.  Sanity&lt;br /&gt;
reproduced 11 NULL &amp;lt;code&amp;gt;FirstTestDate&amp;lt;/code&amp;gt; rows, 12 NULL &amp;lt;code&amp;gt;LastTestDate&amp;lt;/code&amp;gt; rows, and&lt;br /&gt;
confirmed the resulting 164-row projection contains zero &amp;lt;code&amp;gt;unk&amp;lt;/code&amp;gt; rows and&lt;br /&gt;
exactly 10 &amp;lt;code&amp;gt;BoutNumber = 3&amp;lt;/code&amp;gt; rows with no corresponding loaded&lt;br /&gt;
&amp;lt;code&amp;gt;BoutNumber = 2&amp;lt;/code&amp;gt; row, including &amp;lt;code&amp;gt;CH100&amp;lt;/code&amp;gt;, whose eligible negative and&lt;br /&gt;
positive bouts are retained.&lt;br /&gt;
&lt;br /&gt;
== * (#235) SIV_STATUS_BOUT AMA incorrectly claims biography support ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The source row for &amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt; has &amp;lt;code&amp;gt;SSB_in_bio_table = &amp;#039;1&amp;#039;&amp;lt;/code&amp;gt;, but no exact &amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt;&lt;br /&gt;
identifier exists in &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt;.  &amp;lt;code&amp;gt;AME&amp;lt;/code&amp;gt; exists and the source note says&lt;br /&gt;
the row includes &amp;lt;code&amp;gt;AME&amp;lt;/code&amp;gt; samples, but that is not sufficient evidence to replace&lt;br /&gt;
the source identifier.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT source.*&lt;br /&gt;
  FROM clean.siv_status_bout AS source&lt;br /&gt;
  WHERE source.ssb_b_animid = &amp;#039;AMA&amp;#039;&lt;br /&gt;
    AND source.ssb_in_bio_table = &amp;#039;1&amp;#039;&lt;br /&gt;
    AND NOT EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
            FROM sokwedb.biography_data AS biography&lt;br /&gt;
            WHERE biography.animid = source.ssb_b_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns one &amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt; negative bout from 2014-03-20 through&lt;br /&gt;
2019-04-05, with tests from 2017-01-31 through 2019-04-05 and bout number 1.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: exclude this exact complete source&lt;br /&gt;
tuple.  Do not map &amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;AME&amp;lt;/code&amp;gt;, recompute the flag, or create biography&lt;br /&gt;
support.  Preserve all other unsupported source identifiers whose flags are&lt;br /&gt;
&amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, as the target intentionally permits identifiers outside biography.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  The projection excludes the exact&lt;br /&gt;
&amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt; tuple, including its literal doubled-double-quote note text.  Sanity&lt;br /&gt;
confirmed the 21 remaining unsupported identifiers (&amp;lt;code&amp;gt;CH064&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;CH128&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;TUM&amp;lt;/code&amp;gt;) are retained with &amp;lt;code&amp;gt;InBioTable = FALSE&amp;lt;/code&amp;gt; and that no eligible row&lt;br /&gt;
has an &amp;lt;code&amp;gt;InBioTable&amp;lt;/code&amp;gt;/biography-support mismatch.&lt;br /&gt;
&lt;br /&gt;
== (#236) SIV_STATUS_BOUT absent notes cannot satisfy the target ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The source contains 169 NULL &amp;lt;code&amp;gt;SSB_notes&amp;lt;/code&amp;gt; values, while target &amp;lt;code&amp;gt;Notes&amp;lt;/code&amp;gt; is&lt;br /&gt;
non-NULL.  The target&amp;#039;s &amp;lt;code&amp;gt;notonlyspaces_check&amp;lt;/code&amp;gt; explicitly permits the empty&lt;br /&gt;
string as the representation of no note.&lt;br /&gt;
&lt;br /&gt;
After the approved Problems #233 through #235 exclusion union, 158 eligible&lt;br /&gt;
rows have NULL notes and six have non-NULL notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COUNT(*)&lt;br /&gt;
  FROM clean.siv_status_bout AS source&lt;br /&gt;
  WHERE source.ssb_notes IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns 169.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: map an eligible NULL note to the&lt;br /&gt;
empty string with &amp;lt;code&amp;gt;COALESCE(source.ssb_notes, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt;.  Preserve every non-NULL&lt;br /&gt;
note exactly.  Sanity and validation must prove that this is the only notes&lt;br /&gt;
transformation and that no projected note contains only spaces.&lt;br /&gt;
&lt;br /&gt;
Problems #233, #234, and #235 contain 5, 12, and 1 rows respectively.  One&lt;br /&gt;
&amp;lt;code&amp;gt;TT&amp;lt;/code&amp;gt; row overlaps #233 and #234, so their union excludes 17 distinct rows from&lt;br /&gt;
181 and leaves 164 rows for conversion.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  &amp;lt;code&amp;gt;COALESCE(source.ssb_notes, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt; is&lt;br /&gt;
the projection&amp;#039;s only notes transformation.  Read-only sanity and validation&lt;br /&gt;
against the local PostgreSQL 18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; database confirmed 158 loaded&lt;br /&gt;
empty notes came only from NULL source notes, all 6 non-NULL notes are&lt;br /&gt;
preserved byte-for-byte, and no loaded note contains only spaces.&lt;br /&gt;
&lt;br /&gt;
The direct set-based loader&lt;br /&gt;
(&amp;lt;code&amp;gt;conversion/load_siv_status_bout.sql&amp;lt;/code&amp;gt;) inserted exactly 164 rows: 132 &amp;lt;code&amp;gt;neg&amp;lt;/code&amp;gt;,&lt;br /&gt;
32 &amp;lt;code&amp;gt;pos&amp;lt;/code&amp;gt;, 0 &amp;lt;code&amp;gt;unk&amp;lt;/code&amp;gt;; 154 &amp;lt;code&amp;gt;BoutNumber = 1&amp;lt;/code&amp;gt;, 0 &amp;lt;code&amp;gt;BoutNumber = 2&amp;lt;/code&amp;gt;, 10&lt;br /&gt;
&amp;lt;code&amp;gt;BoutNumber = 3&amp;lt;/code&amp;gt;; 142 &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt; and 22 &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;InBioTable&amp;lt;/code&amp;gt; values; and 158&lt;br /&gt;
empty/6 nonempty notes.  Bidirectional &amp;lt;code&amp;gt;EXCEPT ALL&amp;lt;/code&amp;gt; between the canonical&lt;br /&gt;
projection and the loaded rows returned no differences in either direction.&lt;br /&gt;
An intentional post-insert failure (&amp;lt;code&amp;gt;SELECT 1/0&amp;lt;/code&amp;gt;) rolled the single&lt;br /&gt;
transaction back to zero target rows.  An authorized &amp;lt;code&amp;gt;TRUNCATE&amp;lt;/code&amp;gt; of only the&lt;br /&gt;
disposable target table followed by an unchanged sanity/load/validation&lt;br /&gt;
sequence reproduced the identical sorted ten-column payload digest,&lt;br /&gt;
&amp;lt;code&amp;gt;8b243c9b7f53d79f9fa1eae640e26b51&amp;lt;/code&amp;gt;, on both runs.  &amp;lt;code&amp;gt;git diff --check&amp;lt;/code&amp;gt;, a&lt;br /&gt;
local-only Make dry run of the full &amp;lt;code&amp;gt;load_data&amp;lt;/code&amp;gt; sequence, and a direct &amp;lt;code&amp;gt;m4&amp;lt;/code&amp;gt;&lt;br /&gt;
and repository doc-build render of &amp;lt;code&amp;gt;doc/src/housekeeping/siv_status_bout.m4&amp;lt;/code&amp;gt;&lt;br /&gt;
all passed without error.&lt;br /&gt;
&lt;br /&gt;
== (#237) SUBADULT_ARRIVALS_LOG FirstTikiDate is declared NOT NULL contrary to its own documentation ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;db/schemas/housekeeping/tables/create/subadult_arrivals_log.m4&amp;lt;/code&amp;gt; declares&lt;br /&gt;
&amp;lt;code&amp;gt;firsttikidate DATE NOT NULL&amp;lt;/code&amp;gt;, but the owning documentation, added in the&lt;br /&gt;
same commit, says the column &amp;quot;may be NULL when there is no information on&lt;br /&gt;
the date of permanent addition.&amp;quot;  156 of 209 source rows (75%) have a NULL&lt;br /&gt;
&amp;lt;code&amp;gt;SA_first_tiki_date&amp;lt;/code&amp;gt;, each carrying an explanatory note describing why no&lt;br /&gt;
discrete tiki date applies.  A NULL &amp;lt;code&amp;gt;FirstTikiDate&amp;lt;/code&amp;gt; is this table&amp;#039;s normal,&lt;br /&gt;
permanent representation for most individuals, not missing data awaiting&lt;br /&gt;
entry.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT count(*) FILTER (WHERE sa_first_tiki_date IS NULL) AS null_date&lt;br /&gt;
     , count(*) FILTER (WHERE sa_first_tiki_date IS NOT NULL) AS dated&lt;br /&gt;
  FROM clean.subadult_arrivals_log;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns 156 NULL and 53 non-NULL rows.  &amp;lt;code&amp;gt;git show cc3640b&lt;br /&gt;
--stat&amp;lt;/code&amp;gt; shows the table, its documentation, its indexes, and its trigger&lt;br /&gt;
were all added in one commit, so this contradiction has existed since the&lt;br /&gt;
table was created and was never exercised by a loader.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: correct the schema defect&lt;br /&gt;
rather than exclude the 156 legitimately NULL-dated rows.  Remove &amp;lt;code&amp;gt;NOT&lt;br /&gt;
NULL&amp;lt;/code&amp;gt; from the &amp;lt;code&amp;gt;firsttikidate&amp;lt;/code&amp;gt; column in the owning create M4, regenerate&lt;br /&gt;
the owned table SQL artifact, and verify the corrected nullability in a&lt;br /&gt;
disposable or migrated database before loading.  Leave the unique &amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt;&lt;br /&gt;
index, the &amp;lt;code&amp;gt;FirstTikiDate&amp;lt;/code&amp;gt; index, the &amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt; foreign key and &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;&lt;br /&gt;
check, and the update trigger unchanged.  This is a target-schema&lt;br /&gt;
correction, not a clean-stage mutation, and excludes no source rows on its&lt;br /&gt;
own.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  &amp;lt;code&amp;gt;NOT NULL&amp;lt;/code&amp;gt; was removed from&lt;br /&gt;
&amp;lt;code&amp;gt;firsttikidate&amp;lt;/code&amp;gt; in&lt;br /&gt;
&amp;lt;code&amp;gt;db/schemas/housekeeping/tables/create/subadult_arrivals_log.m4&amp;lt;/code&amp;gt;; the owned&lt;br /&gt;
table SQL artifact was regenerated with &amp;lt;code&amp;gt;make db/schemas/createtables.sql&amp;lt;/code&amp;gt;&lt;br /&gt;
and shows &amp;lt;code&amp;gt;firsttikidate DATE&amp;lt;/code&amp;gt; with no other column or constraint changed.&lt;br /&gt;
&amp;lt;code&amp;gt;ALTER TABLE housekeeping.subadult_arrivals_log ALTER COLUMN firsttikidate&lt;br /&gt;
DROP NOT NULL&amp;lt;/code&amp;gt; was applied to the local PostgreSQL 18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; disposable&lt;br /&gt;
database, and &amp;lt;code&amp;gt;information_schema.columns&amp;lt;/code&amp;gt; confirmed &amp;lt;code&amp;gt;firsttikidate&amp;lt;/code&amp;gt; is&lt;br /&gt;
nullable while &amp;lt;code&amp;gt;id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;animid&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;notes&amp;lt;/code&amp;gt; remain &amp;lt;code&amp;gt;NOT NULL&amp;lt;/code&amp;gt; and the unique&lt;br /&gt;
&amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt; index, &amp;lt;code&amp;gt;FirstTikiDate&amp;lt;/code&amp;gt; index, &amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt; foreign key and &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;&lt;br /&gt;
check, and update trigger are unchanged.&lt;br /&gt;
&lt;br /&gt;
== * (#238) SUBADULT_ARRIVALS_LOG CT&amp;#039;s FirstTikiDate follows CT&amp;#039;s own departure date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;CT&amp;lt;/code&amp;gt;&amp;#039;s parsed &amp;lt;code&amp;gt;FirstTikiDate&amp;lt;/code&amp;gt; (1986-03-01) is after &amp;lt;code&amp;gt;CT&amp;lt;/code&amp;gt;&amp;#039;s own&lt;br /&gt;
&amp;lt;code&amp;gt;BIOGRAPHY_DATA.DepartDate&amp;lt;/code&amp;gt; (1985-09-30).  The target has no constraint&lt;br /&gt;
tying &amp;lt;code&amp;gt;FirstTikiDate&amp;lt;/code&amp;gt; to biography dates, so this row would load without&lt;br /&gt;
error, but a tiki date recorded after an individual&amp;#039;s departure is not&lt;br /&gt;
trustworthy evidence of when that individual was added to the tiki sheets.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT source.*, biography.departdate&lt;br /&gt;
  FROM clean.subadult_arrivals_log AS source&lt;br /&gt;
  JOIN sokwedb.biography_data AS biography&lt;br /&gt;
    ON biography.animid = source.sa_b_animid&lt;br /&gt;
  WHERE to_date(source.sa_first_tiki_date, &amp;#039;MM/DD/YY&amp;#039;) &amp;gt; biography.departdate;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns exactly one row: &amp;lt;code&amp;gt;CT&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;03/01/86&amp;lt;/code&amp;gt;, with note &amp;quot;Btwn&lt;br /&gt;
DOB and CA death, arrivals same as CA; Elo confirmed that always accounted&lt;br /&gt;
for on tikis afterwards.&amp;quot;  No other row falls outside its individual&amp;#039;s&lt;br /&gt;
birth/entry/departure window.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: exclude this exact row rather&lt;br /&gt;
than preserve it as a documented semantic anomaly.  Match on the exact&lt;br /&gt;
source tuple (&amp;lt;code&amp;gt;sa_b_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;sa_first_tiki_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;sa_notes&amp;lt;/code&amp;gt;) with&lt;br /&gt;
NULL-safe equality; do not exclude by &amp;lt;code&amp;gt;AnimID&amp;lt;/code&amp;gt; alone or use a broad&lt;br /&gt;
date-comparison predicate as the operative exclusion, and do not infer a&lt;br /&gt;
corrected date.  This is the only approved exclusion; it leaves 208&lt;br /&gt;
eligible rows.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/subadult_arrivals_log_projection.sql&amp;lt;/code&amp;gt; encodes the exact&lt;br /&gt;
one-row &amp;lt;code&amp;gt;CT&amp;lt;/code&amp;gt; tuple.  Read-only sanity, run against the local PostgreSQL&lt;br /&gt;
18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; database, reproduced this exact tuple with bidirectional&lt;br /&gt;
&amp;lt;code&amp;gt;EXCEPT ALL&amp;lt;/code&amp;gt; and confirmed no other eligible row falls outside its&lt;br /&gt;
individual&amp;#039;s birth/entry/departure window.&lt;br /&gt;
&lt;br /&gt;
== (#239) SUBADULT_ARRIVALS_LOG absent notes cannot satisfy the target ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The source contains 6 NULL &amp;lt;code&amp;gt;SA_notes&amp;lt;/code&amp;gt; values, all outside the Problem #238&lt;br /&gt;
exclusion.  The target &amp;lt;code&amp;gt;Notes&amp;lt;/code&amp;gt; is NOT NULL but explicitly permits the empty&lt;br /&gt;
string.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT count(*)&lt;br /&gt;
  FROM clean.subadult_arrivals_log&lt;br /&gt;
  WHERE sa_notes IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns 6.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: map an eligible NULL note to&lt;br /&gt;
the empty string with &amp;lt;code&amp;gt;COALESCE(source.sa_notes, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt;.  Preserve every&lt;br /&gt;
non-NULL note exactly, including two rows (&amp;lt;code&amp;gt;FAD&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;KEA&amp;lt;/code&amp;gt;) that carry leading&lt;br /&gt;
or trailing spaces around substantive text.  Sanity and validation must&lt;br /&gt;
prove that this is the only notes transformation and that no projected&lt;br /&gt;
note contains only spaces.  This mapping, together with Problem #237&amp;#039;s&lt;br /&gt;
schema fix and Problem #238&amp;#039;s exclusion, leaves 208 rows for conversion:&lt;br /&gt;
209 source rows minus the one Problem #238 exclusion.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.&lt;br /&gt;
&amp;lt;code&amp;gt;COALESCE(source.sa_notes, &amp;#039;&amp;#039;)&amp;lt;/code&amp;gt; is the projection&amp;#039;s only notes&lt;br /&gt;
transformation. Read-only sanity and validation against the local&lt;br /&gt;
PostgreSQL 18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; database confirmed 6 loaded empty notes came&lt;br /&gt;
only from NULL source notes, all 202 non-NULL notes are preserved&lt;br /&gt;
byte-for-byte (including the two rows with leading/trailing spaces), and&lt;br /&gt;
no loaded note contains only spaces.&lt;br /&gt;
&lt;br /&gt;
The direct set-based loader (&amp;lt;code&amp;gt;conversion/load_subadult_arrivals_log.sql&amp;lt;/code&amp;gt;)&lt;br /&gt;
inserted exactly 208 rows: 156 NULL and 52 non-NULL &amp;lt;code&amp;gt;FirstTikiDate&amp;lt;/code&amp;gt;&lt;br /&gt;
values (extent 1976-05-01 through 2013-09-22, unchanged by the &amp;lt;code&amp;gt;CT&amp;lt;/code&amp;gt;&lt;br /&gt;
exclusion), and 6 empty/202 nonempty notes. Every non-NULL &amp;lt;code&amp;gt;FirstTikiDate&amp;lt;/code&amp;gt;&lt;br /&gt;
was proven to round-trip through &amp;lt;code&amp;gt;to_char(..., &amp;#039;MM/DD/YY&amp;#039;)&amp;lt;/code&amp;gt; back to its&lt;br /&gt;
exact source text. Bidirectional &amp;lt;code&amp;gt;EXCEPT ALL&amp;lt;/code&amp;gt; between the canonical&lt;br /&gt;
projection and the loaded rows returned no differences in either&lt;br /&gt;
direction. An intentional post-insert failure (&amp;lt;code&amp;gt;SELECT 1/0&amp;lt;/code&amp;gt;) rolled the&lt;br /&gt;
single transaction back to zero target rows. An authorized &amp;lt;code&amp;gt;TRUNCATE&amp;lt;/code&amp;gt; of&lt;br /&gt;
only the disposable target table followed by an unchanged&lt;br /&gt;
sanity/load/validation sequence reproduced the identical sorted&lt;br /&gt;
four-column payload digest, &amp;lt;code&amp;gt;6eac68b10ebb802e3c78b9109e28d3d7&amp;lt;/code&amp;gt;, on both&lt;br /&gt;
runs. &amp;lt;code&amp;gt;git diff --check&amp;lt;/code&amp;gt;, a local-only Make dry run of the full&lt;br /&gt;
&amp;lt;code&amp;gt;load_data&amp;lt;/code&amp;gt; sequence, and a direct &amp;lt;code&amp;gt;m4&amp;lt;/code&amp;gt; and repository doc-build render&lt;br /&gt;
of &amp;lt;code&amp;gt;doc/src/housekeeping/subadult_arrivals_log.m4&amp;lt;/code&amp;gt; all passed without&lt;br /&gt;
error.&lt;br /&gt;
&lt;br /&gt;
== (#240) TLK_BRECORD_NOTES_CODES two Abbreviation values carry a trailing space ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 94&amp;#039;s &amp;lt;code&amp;gt;&amp;quot;Abbreviation&amp;quot;&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;l &amp;lt;/code&amp;gt; (trailing space) with &amp;lt;code&amp;gt;&amp;quot;Meaning&amp;quot;&amp;lt;/code&amp;gt; of&lt;br /&gt;
&amp;lt;code&amp;gt;laugh&amp;lt;/code&amp;gt;; &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 186&amp;#039;s &amp;lt;code&amp;gt;&amp;quot;Abbreviation&amp;quot;&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;th &amp;lt;/code&amp;gt; (trailing space) with&lt;br /&gt;
&amp;lt;code&amp;gt;&amp;quot;Meaning&amp;quot;&amp;lt;/code&amp;gt; of &amp;lt;code&amp;gt;threaten&amp;lt;/code&amp;gt;.  Both violate the target&amp;#039;s&lt;br /&gt;
&amp;lt;code&amp;gt;trimmedofspaces_check&amp;lt;/code&amp;gt;.  Trimming either value makes it collide with an&lt;br /&gt;
existing unpadded row (&amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 93 &amp;lt;code&amp;gt;l&amp;lt;/code&amp;gt; -&amp;gt; &amp;lt;code&amp;gt;locomotion/travel&amp;lt;/code&amp;gt;; &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 185&lt;br /&gt;
&amp;lt;code&amp;gt;th&amp;lt;/code&amp;gt; -&amp;gt; &amp;lt;code&amp;gt;throw&amp;lt;/code&amp;gt;), folding directly into the Problem #241 duplicate-&lt;br /&gt;
abbreviation resolution.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;ID&amp;quot;, &amp;quot;Abbreviation&amp;quot;, &amp;quot;Meaning&amp;quot;&lt;br /&gt;
  FROM clean.tlk_brecord_notes_codes&lt;br /&gt;
  WHERE &amp;quot;Abbreviation&amp;quot; IS DISTINCT FROM BTRIM(&amp;quot;Abbreviation&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns exactly two rows: &amp;lt;code&amp;gt;94, &amp;quot;l &amp;quot;, laugh&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;186,&lt;br /&gt;
&amp;quot;th &amp;quot;, threaten&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19, as an explicit exception to&lt;br /&gt;
this repository&amp;#039;s usual prohibition on clean-stage mutation: trim&lt;br /&gt;
&amp;lt;code&amp;gt;&amp;quot;Abbreviation&amp;quot;&amp;lt;/code&amp;gt; in &amp;lt;code&amp;gt;conversion/clean.sql&amp;lt;/code&amp;gt;, in its own reviewable&lt;br /&gt;
statement, rather than in the loader or projection.  An initial&lt;br /&gt;
implementation matched the &amp;lt;code&amp;gt;UPDATE&amp;lt;/code&amp;gt; by exact &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot; IN (94, 186)&amp;lt;/code&amp;gt;; on&lt;br /&gt;
review the investigator preferred applying &amp;lt;code&amp;gt;BTRIM&amp;lt;/code&amp;gt; unconditionally, as a&lt;br /&gt;
general clean-stage normalization, since it produces the identical result&lt;br /&gt;
here (no other row is padded) and reads more simply as ordinary cleanup&lt;br /&gt;
rather than a two-row special case.  The dated source affects exactly the&lt;br /&gt;
same two rows either way, &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 94 and &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 186.  After this fix, &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt;&lt;br /&gt;
94&amp;#039;s abbreviation is &amp;lt;code&amp;gt;l&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 186&amp;#039;s is &amp;lt;code&amp;gt;th&amp;lt;/code&amp;gt;, each now a member of an&lt;br /&gt;
ordinary Problem #241 duplicate group.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  Added an unconditional&lt;br /&gt;
&amp;lt;code&amp;gt;UPDATE ... SET &amp;quot;Abbreviation&amp;quot; = BTRIM(&amp;quot;Abbreviation&amp;quot;)&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/clean.sql&amp;lt;/code&amp;gt;, independently reviewable from the loader, with a&lt;br /&gt;
comment documenting that the dated source&amp;#039;s only two affected rows are&lt;br /&gt;
&amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 94 and &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; 186.  Read-only re-checks against the local&lt;br /&gt;
PostgreSQL 18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; database confirmed the two exact pre-fix rows&lt;br /&gt;
before the trim, that applying the statement leaves zero padded rows&lt;br /&gt;
afterward, and that the two documented rows&amp;#039; post-fix unpadded values&lt;br /&gt;
(&amp;lt;code&amp;gt;l&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;th&amp;lt;/code&amp;gt;) are unchanged from what the narrowly-targeted form would have&lt;br /&gt;
produced.&lt;br /&gt;
&lt;br /&gt;
== * (#241) TLK_BRECORD_NOTES_CODES duplicate abbreviations after trimming ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After Problem #240&amp;#039;s fix, 26 case-insensitive abbreviations occur more than&lt;br /&gt;
once, with genuinely different &amp;lt;code&amp;gt;Meaning&amp;lt;/code&amp;gt; text recorded for the same&lt;br /&gt;
abbreviation depending on note-taking context (for example &amp;lt;code&amp;gt;r&amp;lt;/code&amp;gt; means&lt;br /&gt;
&amp;lt;code&amp;gt;rest&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;reassurance&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;reunion&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;resume&amp;lt;/code&amp;gt; across four separate rows).&lt;br /&gt;
The target permits only one row per case-insensitive abbreviation via a&lt;br /&gt;
unique index on &amp;lt;code&amp;gt;lower(normalize(Abbreviation))&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT lower(normalize(BTRIM(&amp;quot;Abbreviation&amp;quot;))) AS norm_abbr, count(*)&lt;br /&gt;
     , array_agg(&amp;quot;ID&amp;quot; ORDER BY &amp;quot;ID&amp;quot;) AS ids&lt;br /&gt;
     , array_agg(&amp;quot;Meaning&amp;quot; ORDER BY &amp;quot;ID&amp;quot;) AS meanings&lt;br /&gt;
  FROM clean.tlk_brecord_notes_codes&lt;br /&gt;
  GROUP BY 1&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The dated source returns exactly 26 groups covering 55 rows: 24 groups of 2&lt;br /&gt;
rows, one group of 3 (&amp;lt;code&amp;gt;ds&amp;lt;/code&amp;gt;), and one group of 4 (&amp;lt;code&amp;gt;r&amp;lt;/code&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Superseded on 2026-09-19.  An initial approval picked the lowest source&lt;br /&gt;
&amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; in each group to survive.  On review the investigator rejected that&lt;br /&gt;
tie-break: choosing a survivor&amp;#039;s &amp;lt;code&amp;gt;Meaning&amp;lt;/code&amp;gt; over its siblings&amp;#039; is itself an&lt;br /&gt;
uncorroborated judgment call about which recorded meaning is &amp;quot;correct&amp;quot;,&lt;br /&gt;
which is exactly what a targeted, documented exclusion is meant to avoid.&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19 (final): exclude every row in&lt;br /&gt;
each of the 26 case-insensitive duplicate-abbreviation groups, not only the&lt;br /&gt;
non-survivors.  None of the 26 ambiguous abbreviations is loaded under any&lt;br /&gt;
&amp;lt;code&amp;gt;Meaning&amp;lt;/code&amp;gt;; only the 155 rows whose abbreviation is unique in the source&lt;br /&gt;
(after Problem #240&amp;#039;s trim) are eligible.  This drops all 55 grouped rows,&lt;br /&gt;
not 29.  The excluded rows, and their meanings, remain visible, unmodified,&lt;br /&gt;
in &amp;lt;code&amp;gt;clean.tlk_brecord_notes_codes&amp;lt;/code&amp;gt;.  Do not infer, rank, or merge meanings;&lt;br /&gt;
do not pick a survivor by &amp;lt;code&amp;gt;&amp;quot;ID&amp;quot;&amp;lt;/code&amp;gt; or any other criterion.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-19.  &amp;lt;code&amp;gt;conversion/tlk_brecord_notes_codes_projection.sql&amp;lt;/code&amp;gt;&lt;br /&gt;
excludes every row whose case-insensitive abbreviation occurs more than&lt;br /&gt;
once in &amp;lt;code&amp;gt;clean.tlk_brecord_notes_codes&amp;lt;/code&amp;gt; (after the Problem #240 trim),&lt;br /&gt;
proved against an independently-derived 55-row, 26-abbreviation exclusion&lt;br /&gt;
set.  Read-only sanity (&amp;lt;code&amp;gt;conversion/load_tlk_brecord_notes_codes_sanity.sql&amp;lt;/code&amp;gt;)&lt;br /&gt;
against the local PostgreSQL 18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; database confirmed the&lt;br /&gt;
exclusion set matches exactly and that the canonical projection is exactly&lt;br /&gt;
155 rows with 155 distinct case-insensitive abbreviations, none of them one&lt;br /&gt;
of the 26 excluded abbreviations.&lt;br /&gt;
&lt;br /&gt;
The direct set-based loader (&amp;lt;code&amp;gt;conversion/load_tlk_brecord_notes_codes.sql&amp;lt;/code&amp;gt;)&lt;br /&gt;
inserted exactly 155 rows.  Bidirectional &amp;lt;code&amp;gt;EXCEPT ALL&amp;lt;/code&amp;gt; between the&lt;br /&gt;
canonical projection and the loaded rows&lt;br /&gt;
(&amp;lt;code&amp;gt;conversion/load_tlk_brecord_notes_codes_validation.sql&amp;lt;/code&amp;gt;) returned no&lt;br /&gt;
differences in either direction, every loaded abbreviation was proven&lt;br /&gt;
unique under &amp;lt;code&amp;gt;lower(normalize(...))&amp;lt;/code&amp;gt; and to have exactly one matching&lt;br /&gt;
source row, every loaded &amp;lt;code&amp;gt;Meaning&amp;lt;/code&amp;gt; was proven to match that single source&lt;br /&gt;
row exactly, and none of the 26 excluded abbreviations was proven present&lt;br /&gt;
in the loaded rows.  An intentional post-insert failure (&amp;lt;code&amp;gt;SELECT 1/0&amp;lt;/code&amp;gt;)&lt;br /&gt;
rolled the single transaction back to zero target rows.  An authorized&lt;br /&gt;
&amp;lt;code&amp;gt;TRUNCATE&amp;lt;/code&amp;gt; of only the disposable target table followed by an unchanged&lt;br /&gt;
sanity/load/validation sequence, run twice, reproduced the identical&lt;br /&gt;
sorted two-column payload digest, &amp;lt;code&amp;gt;27da98a74e6378218418652dda914c95&amp;lt;/code&amp;gt;, on&lt;br /&gt;
both runs.  &amp;lt;code&amp;gt;git diff --check&amp;lt;/code&amp;gt;, a&lt;br /&gt;
local-only Make dry run of the full &amp;lt;code&amp;gt;load_data&amp;lt;/code&amp;gt; sequence confirming correct&lt;br /&gt;
ordering, a direct Make-driven execution of the new&lt;br /&gt;
&amp;lt;code&amp;gt;load_tlk_brecord_notes_codes_sanity&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;load_tlk_brecord_notes_codes&amp;lt;/code&amp;gt;/&lt;br /&gt;
&amp;lt;code&amp;gt;load_tlk_brecord_notes_codes_validation&amp;lt;/code&amp;gt; targets, and a direct &amp;lt;code&amp;gt;m4&amp;lt;/code&amp;gt; render&lt;br /&gt;
of &amp;lt;code&amp;gt;doc/src/housekeeping/tlk_brecord_notes_codes.m4&amp;lt;/code&amp;gt; all passed without&lt;br /&gt;
error.&lt;br /&gt;
&lt;br /&gt;
== * (#242) FOOD_VARIATIONS eligible rows lack a LocalName ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Problem #161 deferred all of &amp;lt;code&amp;gt;clean.food_variations_lookup&amp;lt;/code&amp;gt; to a future&lt;br /&gt;
conversion because of missing and unresolved &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt; values.  Picking&lt;br /&gt;
this conversion back up: the target &amp;lt;code&amp;gt;housekeeping.food_variations.LocalName&amp;lt;/code&amp;gt;&lt;br /&gt;
is &amp;lt;code&amp;gt;NOT NULL&amp;lt;/code&amp;gt;, but 224 of the source&amp;#039;s 1,575 rows have a NULL&lt;br /&gt;
&amp;lt;code&amp;gt;fvl_fl_local_food_name&amp;lt;/code&amp;gt;.  40 of those 224 rows have a&lt;br /&gt;
&amp;lt;code&amp;gt;fvl_food_spelling_variant&amp;lt;/code&amp;gt; that itself exactly matches an existing&lt;br /&gt;
&amp;lt;code&amp;gt;codes.food_names&amp;lt;/code&amp;gt; entry (for example &amp;lt;code&amp;gt;DONGO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;KIFUMUFUMU&amp;lt;/code&amp;gt;), but the&lt;br /&gt;
other 184 do not, so there is no single reliable substitute value across&lt;br /&gt;
all 224 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT count(*) AS null_localname_rows&lt;br /&gt;
     , count(*) FILTER (&lt;br /&gt;
         WHERE EXISTS (&lt;br /&gt;
           SELECT 1 FROM codes.food_names AS fn&lt;br /&gt;
           WHERE lower(normalize(fn.name)) =&lt;br /&gt;
                   lower(normalize(fvl.fvl_food_spelling_variant))))&lt;br /&gt;
         AS variant_matches_food_names&lt;br /&gt;
  FROM clean.food_variations_lookup AS fvl&lt;br /&gt;
  WHERE fvl.fvl_fl_local_food_name IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-19 snapshot returns 224 and 40.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-19: exclude all 224 rows with a&lt;br /&gt;
NULL &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt;, including the 40 whose &amp;lt;code&amp;gt;Variant&amp;lt;/code&amp;gt; happens to match a&lt;br /&gt;
&amp;lt;code&amp;gt;codes.food_names&amp;lt;/code&amp;gt; entry.  Do not infer a &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt; from the &amp;lt;code&amp;gt;Variant&amp;lt;/code&amp;gt;,&lt;br /&gt;
even for those 40 rows.  This is a straightforward domain predicate&lt;br /&gt;
(&amp;lt;code&amp;gt;fvl_fl_local_food_name IS NOT NULL&amp;lt;/code&amp;gt;), not an exact-tuple exclusion list,&lt;br /&gt;
since excluding by this condition cannot ambiguously over- or&lt;br /&gt;
under-exclude the way a duplicate-abbreviation tie-break could.  This&lt;br /&gt;
leaves at most 1,351 eligible rows before any other approved exclusion.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-20.  &amp;lt;code&amp;gt;conversion/food_variations_projection.sql&amp;lt;/code&amp;gt;&lt;br /&gt;
excludes every row via the domain predicate &amp;lt;code&amp;gt;fvl_fl_local_food_name IS&lt;br /&gt;
NOT NULL&amp;lt;/code&amp;gt;.  Read-only sanity (&amp;lt;code&amp;gt;conversion/load_food_variations_sanity.sql&amp;lt;/code&amp;gt;)&lt;br /&gt;
against the local PostgreSQL 18.6 &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt; database confirmed the exact&lt;br /&gt;
224-row NULL-&amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt; count and that the canonical projection is&lt;br /&gt;
exactly 1,351 rows with 1,351 distinct case-insensitive &amp;lt;code&amp;gt;Variant&amp;lt;/code&amp;gt; values.&lt;br /&gt;
The direct set-based loader (&amp;lt;code&amp;gt;conversion/load_food_variations.sql&amp;lt;/code&amp;gt;)&lt;br /&gt;
inserted exactly 1,351 rows.  Bidirectional &amp;lt;code&amp;gt;EXCEPT ALL&amp;lt;/code&amp;gt; between the&lt;br /&gt;
canonical projection and the loaded rows&lt;br /&gt;
(&amp;lt;code&amp;gt;conversion/load_food_variations_validation.sql&amp;lt;/code&amp;gt;) returned no&lt;br /&gt;
differences in either direction, every loaded &amp;lt;code&amp;gt;Variant&amp;lt;/code&amp;gt; was proven unique&lt;br /&gt;
under &amp;lt;code&amp;gt;lower(normalize(...))&amp;lt;/code&amp;gt;, every loaded &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt; was proven to&lt;br /&gt;
match its source row exactly, and no loaded row was proven to have a NULL&lt;br /&gt;
source &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt;.  An intentional post-insert failure (&amp;lt;code&amp;gt;SELECT 1/0&amp;lt;/code&amp;gt;)&lt;br /&gt;
rolled the single transaction back to zero target rows.  An authorized&lt;br /&gt;
&amp;lt;code&amp;gt;TRUNCATE&amp;lt;/code&amp;gt; of only the disposable target table followed by an unchanged&lt;br /&gt;
sanity/load/validation sequence, run twice, reproduced the identical&lt;br /&gt;
sorted two-column payload digest, &amp;lt;code&amp;gt;c510323457dcb801412441eda0ceb09c&amp;lt;/code&amp;gt;, on&lt;br /&gt;
both runs.  &amp;lt;code&amp;gt;git diff --check&amp;lt;/code&amp;gt;, a local-only Make dry run of the full&lt;br /&gt;
&amp;lt;code&amp;gt;load_data&amp;lt;/code&amp;gt; sequence confirming correct ordering, a direct Make-driven&lt;br /&gt;
execution of the new &amp;lt;code&amp;gt;load_food_variations_sanity&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;load_food_variations&amp;lt;/code&amp;gt;/&lt;br /&gt;
&amp;lt;code&amp;gt;load_food_variations_validation&amp;lt;/code&amp;gt; targets, and a direct &amp;lt;code&amp;gt;m4&amp;lt;/code&amp;gt; render of&lt;br /&gt;
&amp;lt;code&amp;gt;doc/src/housekeeping/food_variations.m4&amp;lt;/code&amp;gt; all passed without error.&lt;br /&gt;
&lt;br /&gt;
== (#243) FOOD_VARIATIONS orphaned LocalName values are not cross-checked ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Of the 1,351 rows surviving Problem #242, 38 have a &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt; that does&lt;br /&gt;
not match any &amp;lt;code&amp;gt;clean.food_lookup&amp;lt;/code&amp;gt; entry case-insensitively, and a&lt;br /&gt;
stricter 79 do not match the smaller, already-loaded production&lt;br /&gt;
&amp;lt;code&amp;gt;codes.food_names&amp;lt;/code&amp;gt; table.  &amp;lt;code&amp;gt;housekeeping.food_variations&amp;lt;/code&amp;gt; declares no&lt;br /&gt;
foreign key to either table, so nothing in the schema forces a&lt;br /&gt;
resolution; several of the mismatches form typo clusters pointing at a&lt;br /&gt;
name that is not itself canonical either (for example eight variants —&lt;br /&gt;
&amp;lt;code&amp;gt;MBOBOGOLO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MTOMBOGOLO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MUTBOGORO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MUTOBOGALO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MUTOBOGO&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;MUTOBOLO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MUTOBORGORO&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MUTOLGURO&amp;lt;/code&amp;gt; — all record &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;MTOBOGOLO&amp;lt;/code&amp;gt;, which appears in neither reference table).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT count(*) FILTER (&lt;br /&gt;
         WHERE NOT EXISTS (&lt;br /&gt;
           SELECT 1 FROM clean.food_lookup AS fl&lt;br /&gt;
           WHERE lower(fl.fl_local_food_name) = lower(fvl.fvl_fl_local_food_name)))&lt;br /&gt;
         AS orphan_vs_food_lookup&lt;br /&gt;
     , count(*) FILTER (&lt;br /&gt;
         WHERE NOT EXISTS (&lt;br /&gt;
           SELECT 1 FROM codes.food_names AS fn&lt;br /&gt;
           WHERE lower(normalize(fn.name)) =&lt;br /&gt;
                   lower(normalize(fvl.fvl_fl_local_food_name))))&lt;br /&gt;
         AS orphan_vs_food_names&lt;br /&gt;
  FROM clean.food_variations_lookup AS fvl&lt;br /&gt;
  WHERE fvl.fvl_fl_local_food_name IS NOT NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The 2026-09-20 snapshot returns 38 and 79, unchanged since 2026-09-19.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-20: do not cross-check &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt;&lt;br /&gt;
against &amp;lt;code&amp;gt;clean.food_lookup&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;codes.food_names&amp;lt;/code&amp;gt;, or any other table for&lt;br /&gt;
this conversion.  Load every eligible row&amp;#039;s &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt; as free text&lt;br /&gt;
exactly as recorded, matching, mismatching, or orphaned alike.  Sanity may&lt;br /&gt;
report the 38/79 counts as informational profiling, but neither figure&lt;br /&gt;
excludes any row.  This mapping excludes no additional rows beyond&lt;br /&gt;
Problem #242&amp;#039;s 224.&lt;br /&gt;
&lt;br /&gt;
Implemented and verified on 2026-09-20.  The canonical projection applies&lt;br /&gt;
no cross-check against &amp;lt;code&amp;gt;clean.food_lookup&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;codes.food_names&amp;lt;/code&amp;gt;; sanity&lt;br /&gt;
reports the 38/79 mismatch counts as &amp;lt;code&amp;gt;RAISE NOTICE&amp;lt;/code&amp;gt; profiling only.&lt;br /&gt;
Validation confirmed every loaded &amp;lt;code&amp;gt;LocalName&amp;lt;/code&amp;gt; matches its source row&lt;br /&gt;
exactly regardless of whether it matches either reference table, and the&lt;br /&gt;
1,351-row loaded count matches &amp;lt;code&amp;gt;1,575 source - 224 Problem #242&lt;br /&gt;
exclusions&amp;lt;/code&amp;gt; exactly, confirming no additional row was excluded under this&lt;br /&gt;
resolution.&lt;br /&gt;
&lt;br /&gt;
== (#244) REPRO_STATES cannot represent adolescent and not-seen days ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;clean.female_reproductive_states.repro_state&amp;lt;/code&amp;gt; contains &amp;lt;code&amp;gt;A&amp;lt;/code&amp;gt; for adolescent&lt;br /&gt;
and SQL NULL for not seen, but production &amp;lt;code&amp;gt;REPRO_STATES.State&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NOT NULL&amp;lt;/code&amp;gt;&lt;br /&gt;
and permits only cycling (&amp;lt;code&amp;gt;C&amp;lt;/code&amp;gt;), pregnant (&amp;lt;code&amp;gt;P&amp;lt;/code&amp;gt;), and lactating (&amp;lt;code&amp;gt;L&amp;lt;/code&amp;gt;).  The&lt;br /&gt;
dated 2026-09-19 source has 7,582 &amp;lt;code&amp;gt;A&amp;lt;/code&amp;gt; rows and 218,370 NULL rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT repro_state, count(*)&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  GROUP BY repro_state&lt;br /&gt;
  ORDER BY repro_state NULLS LAST;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: adolescent &amp;lt;code&amp;gt;A&amp;lt;/code&amp;gt;&lt;br /&gt;
was added to the target state domain and &amp;lt;code&amp;gt;REPRO_STATES.State&amp;lt;/code&amp;gt; was made&lt;br /&gt;
nullable.  Map source &amp;lt;code&amp;gt;A&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;C&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;P&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;LA&amp;lt;/code&amp;gt; to target &amp;lt;code&amp;gt;A&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;C&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;P&amp;lt;/code&amp;gt;, and&lt;br /&gt;
&amp;lt;code&amp;gt;L&amp;lt;/code&amp;gt;; preserve source NULL as target NULL.  Do not carry a previous state&lt;br /&gt;
into a not-seen day.  The generated SQL and documentation were regenerated;&lt;br /&gt;
catalog inspection and a rollback-only NULL-State insert verified the change.&lt;br /&gt;
&lt;br /&gt;
== (#245) REPRO_STATES requires sparse support attributes ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production &amp;lt;code&amp;gt;Source&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;EstrusDayCert&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;LactEndCert&amp;lt;/code&amp;gt; are &amp;lt;code&amp;gt;NOT NULL&amp;lt;/code&amp;gt;, but&lt;br /&gt;
the dated source contains 646,257, 626,192, and 646,602 NULL values&lt;br /&gt;
respectively.  The 30 rows having an EstrusDay but NULL EstrusDayCert are&lt;br /&gt;
valid under the approved sparse representation.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT count(*) FILTER (WHERE state_change_dec_source IS NULL) AS null_source&lt;br /&gt;
     , count(*) FILTER (WHERE day_of_est_cert IS NULL) AS null_ed_cert&lt;br /&gt;
     , count(*) FILTER (WHERE lactat_end_cert IS NULL) AS null_le_cert&lt;br /&gt;
     , count(*) FILTER (&lt;br /&gt;
         WHERE day_of_estr IS NOT NULL AND day_of_est_cert IS NULL)&lt;br /&gt;
         AS estrus_day_without_cert&lt;br /&gt;
  FROM clean.female_reproductive_states;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: the three&lt;br /&gt;
production columns were made nullable.  Preserve source NULL as SQL NULL;&lt;br /&gt;
do not carry values forward and do not convert NULL to empty text.  The&lt;br /&gt;
support tables reject empty keys, and an empty key would invent a category&lt;br /&gt;
rather than represent missing information.  Catalog inspection and a&lt;br /&gt;
rollback-only sparse-row insert verified the change.&lt;br /&gt;
&lt;br /&gt;
== (#246) REPRO_STATES support tables are empty ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The four required support tables are empty.  The source has twelve distinct&lt;br /&gt;
trimmed change-source labels, ED certainty values &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt; plus the&lt;br /&gt;
Problem #247 typo, LE certainty values &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt;, and three parity&lt;br /&gt;
values.  The Access dump contains no lookup tables defining them.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT upper(btrim(state_change_dec_source)) AS code, count(*)&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  WHERE state_change_dec_source IS NOT NULL&lt;br /&gt;
  GROUP BY 1&lt;br /&gt;
  ORDER BY 1;&lt;br /&gt;
&lt;br /&gt;
SELECT parity, count(*)&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  GROUP BY parity&lt;br /&gt;
  ORDER BY parity;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: load uppercase&lt;br /&gt;
parity keys &amp;lt;code&amp;gt;IMMATURE&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;NULLIPAROUS&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;PAROUS&amp;lt;/code&amp;gt;.  Normalize populated&lt;br /&gt;
change-source keys with &amp;lt;code&amp;gt;upper(btrim(...))&amp;lt;/code&amp;gt;, retaining the twelve semantic&lt;br /&gt;
labels and expanding their state prefixes in Description; the exact&lt;br /&gt;
key/description pairs are recorded in &amp;lt;code&amp;gt;repro_states_loader_handoff.md&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Load ED keys &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt; with the approved descriptions based on the&lt;br /&gt;
length of the unobserved gap following the last full swelling.  Load LE keys&lt;br /&gt;
&amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt; with the approved minimum/maximum swelling, 14-day, and&lt;br /&gt;
first-cycle criteria.  ED code &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt; has no support row; see Problem #248, which&lt;br /&gt;
now corrects &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt; in the clean schema rather than excluding it, so no&lt;br /&gt;
row ever carries an unsupported code.  No support row represents NULL; see&lt;br /&gt;
Problem #245.  Sanity verifies bidirectional parity of every support key and&lt;br /&gt;
description; the clean retry loaded exactly 12 change sources, four ED&lt;br /&gt;
certainties, four LE certainties, and three parities.&lt;br /&gt;
&lt;br /&gt;
== (#247) FEMALE_REPRODUCTIVE_STATES has estrus certainty 22 ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
One row, SIF on 2004-03-21, has &amp;lt;code&amp;gt;day_of_est_cert = &amp;#039;22&amp;#039;&amp;lt;/code&amp;gt;.  The investigator&lt;br /&gt;
confirmed that this is a typo for &amp;lt;code&amp;gt;2&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  WHERE day_of_est_cert = &amp;#039;22&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: change &amp;lt;code&amp;gt;22&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;2&amp;lt;/code&amp;gt; in the clean schema.  The focused clean-stage execution changed exactly&lt;br /&gt;
one row; sanity verifies that &amp;lt;code&amp;gt;22&amp;lt;/code&amp;gt; is absent and the corrected &amp;lt;code&amp;gt;2&amp;lt;/code&amp;gt; count is&lt;br /&gt;
8,251.&lt;br /&gt;
&lt;br /&gt;
== (#248) FEMALE_REPRODUCTIVE_STATES estrus certainty 5 is unsupported ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The dated source has 30 rows with &amp;lt;code&amp;gt;day_of_est_cert = &amp;#039;5&amp;#039;&amp;lt;/code&amp;gt;.  The investigator&lt;br /&gt;
provided meanings for ED certainty codes &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; through &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt;; code &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt; has no&lt;br /&gt;
supported meaning.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  WHERE day_of_est_cert = &amp;#039;5&amp;#039;&lt;br /&gt;
  ORDER BY id, date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: temporarily&lt;br /&gt;
exclude exactly these 30 rows from the reproductive-state projection.  Code&lt;br /&gt;
&amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt; is absent from &amp;lt;code&amp;gt;ED_CERTAINTIES&amp;lt;/code&amp;gt;.  Sanity verifies both the 30-row class and&lt;br /&gt;
that it overlaps none of Problems #251 or #254.&lt;br /&gt;
&lt;br /&gt;
Amended by the investigator on 2026-09-24: convert &amp;lt;code&amp;gt;day_of_est_cert = &amp;#039;5&amp;#039;&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt; in the clean schema instead of excluding the 30 rows.  No &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt; value is&lt;br /&gt;
retained, no row is excluded, and no support row is added for &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt;; the&lt;br /&gt;
correction is applied alongside the Problem #247 typo correction in&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/clean.sql&amp;lt;/code&amp;gt;.  Sanity now asserts an exact 2,325-row &amp;lt;code&amp;gt;4&amp;lt;/code&amp;gt; count and&lt;br /&gt;
rejects any remaining &amp;lt;code&amp;gt;5&amp;lt;/code&amp;gt; value as a corrected-domain violation.  The&lt;br /&gt;
approved projection grew from 649,495 to 649,525 rows, and the pre-expansion&lt;br /&gt;
eligible/exclusion counts changed accordingly; the refreshed counts are&lt;br /&gt;
verified in &amp;lt;code&amp;gt;conversion/repro_states_profile.sql&amp;lt;/code&amp;gt;,&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/load_repro_states_sanity.sql&amp;lt;/code&amp;gt;, and&lt;br /&gt;
&amp;lt;code&amp;gt;conversion/load_repro_states_validation.sql&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#249) FEMALE_REPRODUCTIVE_STATES swelling bounds have no target ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Source columns &amp;lt;code&amp;gt;min_sw&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;max_sw&amp;lt;/code&amp;gt; have no &amp;lt;code&amp;gt;REPRO_STATES&amp;lt;/code&amp;gt; destination.  The&lt;br /&gt;
dated source has 160,268 rows where both are populated and 486,621 where both&lt;br /&gt;
are NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT min_sw, max_sw, count(*)&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  GROUP BY min_sw, max_sw&lt;br /&gt;
  ORDER BY min_sw NULLS LAST, max_sw NULLS LAST;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: do not map or&lt;br /&gt;
convert these fields because they are generated by &amp;lt;code&amp;gt;build_swelling_states()&amp;lt;/code&amp;gt;.&lt;br /&gt;
They remain unchanged in clean as audit evidence and are absent from the&lt;br /&gt;
shared production projection.  This decision excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#250) FEMALE_REPRODUCTIVE_STATES has two incorrect female IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Female IDs &amp;lt;code&amp;gt;OBE&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;UWE&amp;lt;/code&amp;gt; have no production biography match, affecting&lt;br /&gt;
1,404 and 29 rows respectively.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT id, count(*)&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  WHERE id IN (&amp;#039;OBE&amp;#039;, &amp;#039;UWE&amp;#039;)&lt;br /&gt;
  GROUP BY id&lt;br /&gt;
  ORDER BY id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: correct &amp;lt;code&amp;gt;OBE&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;POR&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;UWE&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt; in the clean schema.  The focused clean-stage&lt;br /&gt;
execution changed exactly 1,433 rows; sanity verifies 1,404 &amp;lt;code&amp;gt;POR&amp;lt;/code&amp;gt; and 29 &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;&lt;br /&gt;
rows, no old labels, and applies study-boundary checks afterward.&lt;br /&gt;
&lt;br /&gt;
== * (#251) FEMALE_REPRODUCTIVE_STATES dates fall outside study intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
After applying Problem #250&amp;#039;s female corrections, 2,145 dated source rows&lt;br /&gt;
fall before EntryDate or after DepartDate.  This replaces the earlier&lt;br /&gt;
matched-ID-only count of 1,175: correcting &amp;lt;code&amp;gt;OBE&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;POR&amp;lt;/code&amp;gt; adds 677&lt;br /&gt;
before-entry and 293 after-departure rows.  The other corrected ID, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;,&lt;br /&gt;
adds none.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH corrected AS (&lt;br /&gt;
  SELECT source.*&lt;br /&gt;
       , CASE source.id&lt;br /&gt;
           WHEN &amp;#039;OBE&amp;#039; THEN &amp;#039;POR&amp;#039;&lt;br /&gt;
           WHEN &amp;#039;UWE&amp;#039; THEN &amp;#039;MGF&amp;#039;&lt;br /&gt;
           ELSE source.id&lt;br /&gt;
         END AS animid&lt;br /&gt;
    FROM clean.female_reproductive_states AS source)&lt;br /&gt;
SELECT corrected.*&lt;br /&gt;
  FROM corrected&lt;br /&gt;
  JOIN sokwedb.biography_data&lt;br /&gt;
    ON biography_data.animid = corrected.animid&lt;br /&gt;
  WHERE corrected.date &amp;lt; biography_data.entrydate&lt;br /&gt;
        OR corrected.date &amp;gt; biography_data.departdate&lt;br /&gt;
  ORDER BY corrected.animid, corrected.date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: temporarily&lt;br /&gt;
exclude every corrected-female/date row outside the inclusive study interval.&lt;br /&gt;
The shared projection does not alter dates, and sanity independently verifies&lt;br /&gt;
exactly 2,145 exclusions and no overlap with Problems #248 or #254.&lt;br /&gt;
&lt;br /&gt;
== (#252) FEMALE_REPRODUCTIVE_STATES compound offspring IDs represent twins ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Five slash-delimited &amp;lt;code&amp;gt;young_kid_id&amp;lt;/code&amp;gt; values represent twins rather than invalid&lt;br /&gt;
single identifiers.  A single YKID foreign key cannot store both offspring.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT young_kid_id, count(*)&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  WHERE young_kid_id LIKE &amp;#039;%/%&amp;#039;&lt;br /&gt;
  GROUP BY young_kid_id&lt;br /&gt;
  ORDER BY young_kid_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The raw dated counts are &amp;lt;code&amp;gt;GAB2/GAB3&amp;lt;/code&amp;gt; 6, &amp;lt;code&amp;gt;GLI/GLD&amp;lt;/code&amp;gt; 2,013, &amp;lt;code&amp;gt;GY/GL&amp;lt;/code&amp;gt; 296,&lt;br /&gt;
&amp;lt;code&amp;gt;SG/CT&amp;lt;/code&amp;gt; 3,472, and &amp;lt;code&amp;gt;SHO/ROO&amp;lt;/code&amp;gt; 368.  After preceding exclusions, 5,546 source&lt;br /&gt;
rows are eligible and expand to 11,092 target rows; Problem #251 removes 609&lt;br /&gt;
of the &amp;lt;code&amp;gt;SG/CT&amp;lt;/code&amp;gt; rows first.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: split each&lt;br /&gt;
eligible compound row into two target rows, one per offspring, copying all&lt;br /&gt;
other source payload before deriving YKAge separately for each offspring.  The&lt;br /&gt;
projection names only the five approved labels and fails sanity on any new&lt;br /&gt;
slash-delimited label.  It expands 5,546 eligible source rows to 11,092 output&lt;br /&gt;
rows.  See Problem #258 for the resulting female/date/YKID invariant.&lt;br /&gt;
&lt;br /&gt;
== (#253) FEMALE_REPRODUCTIVE_STATES offspring AMA is incorrect ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The unmatched offspring label &amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt; occurs on 287 dated source rows.  The&lt;br /&gt;
investigator confirmed that these values should identify &amp;lt;code&amp;gt;AME&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  WHERE young_kid_id = &amp;#039;AMA&amp;#039;&lt;br /&gt;
  ORDER BY id, date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: correct &amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt; to&lt;br /&gt;
&amp;lt;code&amp;gt;AME&amp;lt;/code&amp;gt; in the clean schema.  The focused clean-stage execution changed exactly&lt;br /&gt;
287 rows, and sanity verifies that &amp;lt;code&amp;gt;AMA&amp;lt;/code&amp;gt; is absent and 287 rows use &amp;lt;code&amp;gt;AME&amp;lt;/code&amp;gt;.&lt;br /&gt;
This correction excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== * (#254) FEMALE_REPRODUCTIVE_STATES offspring TTB2 is unsupported ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The unmatched offspring label &amp;lt;code&amp;gt;TTB2&amp;lt;/code&amp;gt; occurs on 546 dated source rows and has&lt;br /&gt;
no approved production biography identity.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.female_reproductive_states&lt;br /&gt;
  WHERE young_kid_id = &amp;#039;TTB2&amp;#039;&lt;br /&gt;
  ORDER BY id, date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: temporarily&lt;br /&gt;
exclude exactly these 546 rows.  The projection creates no biography row and&lt;br /&gt;
infers no replacement offspring.  Sanity verifies the exact class and that it&lt;br /&gt;
overlaps neither Problem #248 nor Problem #251.  Problem #248&amp;#039;s 2026-09-24&lt;br /&gt;
amendment converts its 30 rows to a supported code rather than excluding&lt;br /&gt;
them, so this class no longer shares an exclusion mechanism with Problem&lt;br /&gt;
#248, but the two remain non-overlapping in the source data.&lt;br /&gt;
&lt;br /&gt;
== * (#255) Source youngest-offspring ages disagree with biography ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Of 268,336 originally resolved source rows, 115,527 source&lt;br /&gt;
&amp;lt;code&amp;gt;youngest_kid_age&amp;lt;/code&amp;gt; values differ from &amp;lt;code&amp;gt;Date - BIOGRAPHY_DATA.BirthDate&amp;lt;/code&amp;gt;.&lt;br /&gt;
After approved corrections, preceding exclusions, and twin expansion, every&lt;br /&gt;
retained YKID has a biography BirthDate, but 219 nonsplit rows derive a&lt;br /&gt;
negative age from -41 through -1.  Production YKAge must be nonnegative.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH expanded AS (&lt;br /&gt;
  SELECT source.id, source.date&lt;br /&gt;
       , CASE btrim(part.ykid) WHEN &amp;#039;AMA&amp;#039; THEN &amp;#039;AME&amp;#039;&lt;br /&gt;
                              ELSE btrim(part.ykid) END AS ykid&lt;br /&gt;
    FROM clean.female_reproductive_states AS source&lt;br /&gt;
    CROSS JOIN LATERAL&lt;br /&gt;
         regexp_split_to_table(source.young_kid_id, &amp;#039;/&amp;#039;) AS part(ykid)&lt;br /&gt;
   WHERE source.young_kid_id IS NOT NULL&lt;br /&gt;
         AND source.young_kid_id &amp;amp;lt;&amp;amp;gt; &amp;#039;TTB2&amp;#039;)&lt;br /&gt;
SELECT expanded.*, biography_data.birthdate&lt;br /&gt;
     , expanded.date - biography_data.birthdate AS derived_age&lt;br /&gt;
  FROM expanded&lt;br /&gt;
  JOIN sokwedb.biography_data&lt;br /&gt;
    ON biography_data.animid = expanded.ykid&lt;br /&gt;
  WHERE expanded.date - biography_data.birthdate &amp;amp;lt; 0&lt;br /&gt;
  ORDER BY expanded.id, expanded.date, expanded.ykid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The implementation profile must additionally apply Problems #248, #250, and&lt;br /&gt;
#251 in projection order and assert the final 219-row count independently.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: ignore source&lt;br /&gt;
&amp;lt;code&amp;gt;youngest_kid_age&amp;lt;/code&amp;gt; and derive target YKAge from target Date minus the resolved&lt;br /&gt;
offspring&amp;#039;s biography BirthDate.  The projection excludes the 219 rows whose&lt;br /&gt;
derived age is negative.  Sanity verifies zero unresolved retained offspring,&lt;br /&gt;
and validation checks every loaded age against biography.  The original&lt;br /&gt;
115,527-value disagreement remains in the reproducible profile.&lt;br /&gt;
&lt;br /&gt;
== * (#256) FEMALE_REPRO_HISTORY is outside reproductive-state scope ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;clean.female_repro_history&amp;lt;/code&amp;gt; is related historical data but is neither a&lt;br /&gt;
duplicate nor a drop-in supplement for &amp;lt;code&amp;gt;clean.female_reproductive_states&amp;lt;/code&amp;gt;.&lt;br /&gt;
It has 128,306 rows, including 764 female/date keys absent from the requested&lt;br /&gt;
daily source, and overlapping fields differ extensively.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT count(*) AS rows&lt;br /&gt;
     , count(DISTINCT (id, date)) AS female_date_keys&lt;br /&gt;
  FROM clean.female_repro_history;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-21: do not include&lt;br /&gt;
&amp;lt;code&amp;gt;clean.female_repro_history&amp;lt;/code&amp;gt; in this conversion.  Retain the reconciliation&lt;br /&gt;
queries as profile evidence only.  This decision does not exclude any row from&lt;br /&gt;
the requested source.&lt;br /&gt;
&lt;br /&gt;
== (#257) REPRO_STATES AnimID updates bypass integrity checks ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;repro_states_func()&amp;lt;/code&amp;gt; checks sex only on INSERT and checks study boundaries&lt;br /&gt;
only on INSERT or a Date change.  A rollback-only probe changed only AnimID on&lt;br /&gt;
a valid female row to male &amp;lt;code&amp;gt;AL&amp;lt;/code&amp;gt; on 2004-11-19, after AL&amp;#039;s 1999-01-13&lt;br /&gt;
DepartDate; the UPDATE was accepted.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_get_functiondef(&amp;#039;sokwedb.repro_states_func()&amp;#039;::regprocedure);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The returned function gates the sex check on &amp;lt;code&amp;gt;TG_OP = &amp;#039;INSERT&amp;#039;&amp;lt;/code&amp;gt; and the date&lt;br /&gt;
check on &amp;lt;code&amp;gt;TG_OP = &amp;#039;INSERT&amp;#039; OR NEW.date &amp;amp;lt;&amp;amp;gt; OLD.date&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: the owning M4&lt;br /&gt;
reruns female-sex validation whenever AnimID changes and reruns study-boundary&lt;br /&gt;
validation whenever AnimID or Date changes.  Rollback-only tests verified that&lt;br /&gt;
updates to a male AnimID, to a female outside her study interval, and to a date&lt;br /&gt;
outside the current female&amp;#039;s study interval are rejected.&lt;br /&gt;
&lt;br /&gt;
== (#258) REPRO_STATES lacks the female/date/offspring invariant ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Production permits duplicate reproductive-state rows.  The earlier proposed&lt;br /&gt;
female/date uniqueness rule is too strict because Problem #252 requires one&lt;br /&gt;
row per twin; the intended identity is female, date, and youngest offspring.&lt;br /&gt;
Rows with no youngest offspring must still be unique per female/date.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT animid, date, ykid, count(*)&lt;br /&gt;
  FROM sokwedb.repro_states&lt;br /&gt;
  GROUP BY animid, date, ykid&lt;br /&gt;
  HAVING count(*) &amp;amp;gt; 1&lt;br /&gt;
  ORDER BY animid, date, ykid NULLS FIRST;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A rollback-only probe inserted two otherwise valid rows for AND on&lt;br /&gt;
2004-11-19 and both were accepted.  The dated approved projection has zero&lt;br /&gt;
duplicate &amp;lt;code&amp;gt;(AnimID, Date, YKID)&amp;lt;/code&amp;gt; keys when NULL YKID is treated as equal.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator and implemented on 2026-09-21: a unique index&lt;br /&gt;
was added on &amp;lt;code&amp;gt;(AnimID, Date, YKID) NULLS NOT DISTINCT&amp;lt;/code&amp;gt;, and the table&lt;br /&gt;
documentation describes one row per female/date/youngest-offspring.&lt;br /&gt;
Rollback-only tests verified that duplicate NULL and non-NULL YKID keys are&lt;br /&gt;
rejected while two distinct twin YKIDs on one female/date are accepted.&lt;br /&gt;
&lt;br /&gt;
== (#259) ELO_RANKS_DAILY_KK_F has rows outside GK&amp;#039;s and LB&amp;#039;s recorded community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;clean.elo_ranks_daily___kk_females&amp;lt;/code&amp;gt; maps directly to &amp;lt;code&amp;gt;sokwedb.elo_ranks_daily_kk_f&amp;lt;/code&amp;gt;,&lt;br /&gt;
a table scoped to the &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt; (Kasekela) community.  An inclusive join to&lt;br /&gt;
&amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; found 821 rows for four females that do not fall within a &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt;&lt;br /&gt;
interval on their ranked date:&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;GK&amp;lt;/code&amp;gt;, 181 rows, 1972-11-01 through 1973-04-30: 61 days have no &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt;&lt;br /&gt;
  interval at all, and the remaining 120 days fall within an &amp;lt;code&amp;gt;HK&amp;lt;/code&amp;gt; (Kahama)&lt;br /&gt;
  interval.&lt;br /&gt;
* &amp;lt;code&amp;gt;LB&amp;lt;/code&amp;gt;, 638 rows, 1973-01-01 through 1974-09-30: all 638 days fall within an&lt;br /&gt;
  &amp;lt;code&amp;gt;HK&amp;lt;/code&amp;gt; interval.&lt;br /&gt;
* &amp;lt;code&amp;gt;ML&amp;lt;/code&amp;gt;, 1 row, 1986-10-24: no &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; interval (see Problem #260).&lt;br /&gt;
* &amp;lt;code&amp;gt;AL&amp;lt;/code&amp;gt;, 1 row, 1999-01-14, in &amp;lt;code&amp;gt;sokwedb.elo_ranks_daily_kk_m&amp;lt;/code&amp;gt;: no &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt;&lt;br /&gt;
  interval (see Problem #260).&lt;br /&gt;
&lt;br /&gt;
This spans the historical &amp;lt;code&amp;gt;KK&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;HK&amp;lt;/code&amp;gt; community split (&amp;lt;code&amp;gt;HK&amp;lt;/code&amp;gt; starts 1973-01-01,&lt;br /&gt;
ends 1977-12-31 per &amp;lt;code&amp;gt;codes.COMM_IDS&amp;lt;/code&amp;gt;).  Neither the target schema nor the&lt;br /&gt;
current loader constrains ranking dates to a &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; interval, so nothing&lt;br /&gt;
in production would reject any treatment of these rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT e.date, e.individual&lt;br /&gt;
  FROM clean.elo_ranks_daily___kk_females e&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
    SELECT 1 FROM sokwedb.comm_membs cm&lt;br /&gt;
     WHERE cm.animid = e.individual&lt;br /&gt;
       AND e.date BETWEEN cm.startdate AND cm.enddate&lt;br /&gt;
       AND cm.commid = &amp;#039;KK&amp;#039;&lt;br /&gt;
  )&lt;br /&gt;
 ORDER BY e.date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Returns 820 rows (&amp;lt;code&amp;gt;GK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;LB&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;ML&amp;lt;/code&amp;gt;); the equivalent query against&lt;br /&gt;
&amp;lt;code&amp;gt;clean.elo_ranks_daily___kk_males&amp;lt;/code&amp;gt; returns 1 row (&amp;lt;code&amp;gt;AL&amp;lt;/code&amp;gt;).&lt;br /&gt;
&amp;lt;code&amp;gt;clean.elo_ranks_daily___mt_females&amp;lt;/code&amp;gt; has zero such rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-22: preserve all 821 rows unchanged.&lt;br /&gt;
The daily Elo ranking is a precomputed, published output; the conversion&amp;#039;s&lt;br /&gt;
role is to carry it forward as supplied, not to re-scope it by community&lt;br /&gt;
after the fact.  No row is excluded and no rank is recalculated.  This&lt;br /&gt;
decision applies identically to the 758 days recorded in &amp;lt;code&amp;gt;HK&amp;lt;/code&amp;gt; and the 63 days&lt;br /&gt;
with no &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; interval at all.&lt;br /&gt;
&lt;br /&gt;
== (#260) ELO_RANKS_DAILY ranking dates trail DepartDate by one day for ML and AL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Two rows are dated exactly one day after the individual&amp;#039;s&lt;br /&gt;
&amp;lt;code&amp;gt;BIOGRAPHY_DATA.DepartDate&amp;lt;/code&amp;gt;:&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;ML&amp;lt;/code&amp;gt; in &amp;lt;code&amp;gt;clean.elo_ranks_daily___kk_females&amp;lt;/code&amp;gt;: ranked 1986-10-24, DepartDate&lt;br /&gt;
  1986-10-23.&lt;br /&gt;
* &amp;lt;code&amp;gt;AL&amp;lt;/code&amp;gt; in &amp;lt;code&amp;gt;clean.elo_ranks_daily___kk_males&amp;lt;/code&amp;gt;: ranked 1999-01-14, DepartDate&lt;br /&gt;
  1999-01-13.&lt;br /&gt;
&lt;br /&gt;
Neither row is otherwise malformed, and the target schema does not constrain&lt;br /&gt;
ranking dates to an individual&amp;#039;s study interval, so both would load without&lt;br /&gt;
error.  Both rows are also members of the 63 &amp;quot;no &amp;lt;code&amp;gt;COMM_MEMBS&amp;lt;/code&amp;gt; interval&amp;quot; rows&lt;br /&gt;
described in Problem #259.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT e.date, e.individual, b.departdate&lt;br /&gt;
  FROM clean.elo_ranks_daily___kk_females e&lt;br /&gt;
  JOIN sokwedb.biography_data b ON b.animid = e.individual&lt;br /&gt;
 WHERE e.date = b.departdate + 1;&lt;br /&gt;
-- and the equivalent query against clean.elo_ranks_daily___kk_males&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Returns exactly the &amp;lt;code&amp;gt;ML&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;AL&amp;lt;/code&amp;gt; rows above; no other individual in any of&lt;br /&gt;
the three sources is ranked after their own DepartDate.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-22: preserve both rows unchanged, as&lt;br /&gt;
a final recorded ranking day.  Because each ranked day is a complete &amp;lt;code&amp;gt;1..N&amp;lt;/code&amp;gt;&lt;br /&gt;
permutation, excluding either individual&amp;#039;s single row would leave that day&lt;br /&gt;
looking complete while no longer being one, so no partial-day exclusion is&lt;br /&gt;
created for either row.  This decision was made together with, and is&lt;br /&gt;
consistent with, Problem #259&amp;#039;s community-membership disposition, without&lt;br /&gt;
treating the two as the same underlying question.&lt;br /&gt;
&lt;br /&gt;
== (#261) ELO_RANKS_DAILY documentation misstates the ExpNumBeaten formula ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The shared documentation macro (&amp;lt;code&amp;gt;doc/include/elo_ranks_daily_macro.m4&amp;lt;/code&amp;gt;)&lt;br /&gt;
defines &amp;lt;code&amp;gt;ExpNumBeaten&amp;lt;/code&amp;gt; as &amp;lt;code&amp;gt;EloCardinal&amp;lt;/code&amp;gt; multiplied by the number of&lt;br /&gt;
individuals ranked that day.  Every row in all three sources contradicts that&lt;br /&gt;
formula.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT e.date, e.individual, e.expnumbeaten, e.elocardinal, d.n,&lt;br /&gt;
       e.elocardinal * d.n       AS documented_formula,&lt;br /&gt;
       e.elocardinal * (d.n - 1) AS observed_formula&lt;br /&gt;
  FROM clean.elo_ranks_daily___kk_females e&lt;br /&gt;
  JOIN (SELECT date, count(*) AS n&lt;br /&gt;
          FROM clean.elo_ranks_daily___kk_females&lt;br /&gt;
         GROUP BY date) d ON d.date = e.date&lt;br /&gt;
 ORDER BY abs(e.expnumbeaten - e.elocardinal * d.n) DESC&lt;br /&gt;
 LIMIT 5;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Every row instead matches &amp;lt;code&amp;gt;EloCardinal * (N - 1)&amp;lt;/code&amp;gt;, where &amp;lt;code&amp;gt;N&amp;lt;/code&amp;gt; is the number of&lt;br /&gt;
individuals ranked that day, within 0.00000006 across all three sources --&lt;br /&gt;
i.e. &amp;lt;code&amp;gt;ExpNumBeaten&amp;lt;/code&amp;gt; compares each individual against every other ranked&lt;br /&gt;
individual, not against themselves.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-22: correct&lt;br /&gt;
&amp;lt;code&amp;gt;doc/include/elo_ranks_daily_macro.m4&amp;lt;/code&amp;gt; to describe &amp;lt;code&amp;gt;ExpNumBeaten&amp;lt;/code&amp;gt; as&lt;br /&gt;
&amp;lt;code&amp;gt;EloCardinal * (N - 1)&amp;lt;/code&amp;gt;.  Source &amp;lt;code&amp;gt;ExpNumBeaten&amp;lt;/code&amp;gt; values are copied into&lt;br /&gt;
production unchanged regardless; only the written definition changes, and&lt;br /&gt;
this excludes no rows.&lt;br /&gt;
&lt;br /&gt;
== (#262) Shared ELO_RANKS_DAILY index macros target only KK_F ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The shared &amp;lt;code&amp;gt;elo_ranks_daily_indexes_create&amp;lt;/code&amp;gt; and&lt;br /&gt;
&amp;lt;code&amp;gt;elo_ranks_daily_indexes_drop&amp;lt;/code&amp;gt; macros ignored their sex and community&lt;br /&gt;
arguments and hard-coded every index name and target relation to&lt;br /&gt;
&amp;lt;code&amp;gt;ELO_RANKS_DAILY_KK_F&amp;lt;/code&amp;gt;.  Consequently, KK_F had its documented unique and&lt;br /&gt;
secondary indexes, while KK_M, MT_F, and MT_M had only their primary-key&lt;br /&gt;
indexes.  In particular, production did not enforce one row per&lt;br /&gt;
&amp;lt;code&amp;gt;(Date, Individual)&amp;lt;/code&amp;gt; on those three tables.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT tablename, indexname, indexdef&lt;br /&gt;
  FROM pg_indexes&lt;br /&gt;
 WHERE schemaname = &amp;#039;sokwedb&amp;#039;&lt;br /&gt;
   AND tablename LIKE &amp;#039;elo_ranks_daily_%&amp;#039;&lt;br /&gt;
 ORDER BY tablename, indexname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Before the repair this returned nine indexes for KK_F and only the primary&lt;br /&gt;
key for each of KK_M, MT_F, and MT_M.  A rollback-only probe confirmed that&lt;br /&gt;
KK_M accepted a duplicate &amp;lt;code&amp;gt;(Date, Individual)&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Implemented on 2026-09-23: both owning macros now derive the relation and&lt;br /&gt;
index names from their sex and community arguments.  All four generated&lt;br /&gt;
create/drop pairs were reviewed independently and installed, leaving each&lt;br /&gt;
table with its primary key plus the expected unique and seven secondary&lt;br /&gt;
indexes.  Rollback-only probes verified that duplicate INSERT and UPDATE&lt;br /&gt;
collisions are rejected in KK_F, KK_M, MT_F, and MT_M, and that all probe&lt;br /&gt;
rows are removed by rollback.&lt;br /&gt;
&lt;br /&gt;
== (#263) BIOGRAPHY_DATA reverse sex validation omits three ELO_RANKS_DAILY tables ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Each ELO_RANKS_DAILY table&amp;#039;s own trigger requires its Individual to have the&lt;br /&gt;
sex encoded by the table suffix.  The reverse &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; trigger,&lt;br /&gt;
however, checked only &amp;lt;code&amp;gt;ELO_RANKS_DAILY_KK_F&amp;lt;/code&amp;gt;.  After a valid row was inserted,&lt;br /&gt;
changing the referenced biography sex could therefore leave KK_M, MT_F, or&lt;br /&gt;
MT_M in a state their own insert/update trigger would reject.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
The defect was reproduced with rollback-only female and male biography&lt;br /&gt;
fixtures.  After inserting one valid rank row in turn, the following updates&lt;br /&gt;
were accepted for MT_F, KK_M, and MT_M, while the equivalent KK_F update was&lt;br /&gt;
rejected:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE sokwedb.biography_data&lt;br /&gt;
   SET sex = &amp;#039;M&amp;#039;&lt;br /&gt;
 WHERE animid = &amp;#039;female fixture&amp;#039;;&lt;br /&gt;
&lt;br /&gt;
UPDATE sokwedb.biography_data&lt;br /&gt;
   SET sex = &amp;#039;F&amp;#039;&lt;br /&gt;
 WHERE animid = &amp;#039;male fixture&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Implemented on 2026-09-23: the owning biography trigger M4 now uses one&lt;br /&gt;
parameterized reverse-check fragment for KK_F, KK_M, MT_F, and MT_M.  Focused&lt;br /&gt;
rollback-only tests verified that female-to-male changes are rejected when&lt;br /&gt;
the individual is referenced by KK_F or MT_F, male-to-female changes are&lt;br /&gt;
rejected when referenced by KK_M or MT_M, and no fixture rows remain.&lt;br /&gt;
&lt;br /&gt;
== * (#264) BIOGRAPHY_DATA REPRO_STATES diagnostic can produce a NULL RAISE option ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
While selecting an existing female for the Problem #263 rollback probe,&lt;br /&gt;
changing &amp;lt;code&amp;gt;AND&amp;lt;/code&amp;gt; from female to male reached the intended reverse REPRO_STATES&lt;br /&gt;
integrity check but failed while constructing its diagnostic.  Several&lt;br /&gt;
REPRO_STATES fields included in &amp;lt;code&amp;gt;DETAIL&amp;lt;/code&amp;gt; are nullable and are concatenated&lt;br /&gt;
without &amp;lt;code&amp;gt;textualize()&amp;lt;/code&amp;gt;.  A NULL field makes the entire &amp;lt;code&amp;gt;DETAIL&amp;lt;/code&amp;gt; expression&lt;br /&gt;
NULL, which PL/pgSQL rejects before it can raise the intended&lt;br /&gt;
&amp;lt;code&amp;gt;integrity_constraint_violation&amp;lt;/code&amp;gt; message.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
BEGIN;&lt;br /&gt;
UPDATE sokwedb.biography_data&lt;br /&gt;
   SET sex = &amp;#039;M&amp;#039;&lt;br /&gt;
 WHERE animid = &amp;#039;AND&amp;#039;;&lt;br /&gt;
ROLLBACK;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This rollback-only probe fails with SQLSTATE &amp;lt;code&amp;gt;22004&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;RAISE statement option&lt;br /&gt;
cannot be null&amp;lt;/code&amp;gt;, in &amp;lt;code&amp;gt;lib.biography_data_func()&amp;lt;/code&amp;gt; instead of reporting the&lt;br /&gt;
referencing REPRO_STATES row.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Unresolved and outside the Elo loader scope.  Use &amp;lt;code&amp;gt;textualize()&amp;lt;/code&amp;gt; or otherwise&lt;br /&gt;
NULL-safe formatting for every nullable REPRO_STATES value included in the&lt;br /&gt;
trigger diagnostic.  The Elo reverse-trigger test uses isolated biography&lt;br /&gt;
fixtures so this pre-existing error path does not mask Problem #263.&lt;br /&gt;
&lt;br /&gt;
== * (#265) BIOGRAPHY_DATA ARRIVALS diagnostic references undeclared a_pid ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
While selecting an existing male for the Problem #263 rollback probe,&lt;br /&gt;
changing &amp;lt;code&amp;gt;AL&amp;lt;/code&amp;gt; from male to female reached the ARRIVALS check for a female&lt;br /&gt;
assigned the male swelling code.  The block&amp;#039;s diagnostic references &amp;lt;code&amp;gt;a_pid&amp;lt;/code&amp;gt;,&lt;br /&gt;
but the block neither declares that variable nor selects &amp;lt;code&amp;gt;roles.pid&amp;lt;/code&amp;gt; into it.&lt;br /&gt;
PostgreSQL therefore fails while parsing the intended diagnostic expression.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
BEGIN;&lt;br /&gt;
UPDATE sokwedb.biography_data&lt;br /&gt;
   SET sex = &amp;#039;F&amp;#039;&lt;br /&gt;
 WHERE animid = &amp;#039;AL&amp;#039;;&lt;br /&gt;
ROLLBACK;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This rollback-only probe fails with SQLSTATE &amp;lt;code&amp;gt;42703&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;column &amp;quot;a_pid&amp;quot; does not&lt;br /&gt;
exist&amp;lt;/code&amp;gt;, in &amp;lt;code&amp;gt;lib.biography_data_func()&amp;lt;/code&amp;gt; instead of reporting the referencing&lt;br /&gt;
ARRIVALS row.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Unresolved and outside the Elo loader scope.  Declare &amp;lt;code&amp;gt;a_pid&amp;lt;/code&amp;gt;, select&lt;br /&gt;
&amp;lt;code&amp;gt;roles.pid&amp;lt;/code&amp;gt; into it in each applicable ARRIVALS query, and retain it in the&lt;br /&gt;
diagnostic.  The Elo reverse-trigger test uses isolated biography fixtures so&lt;br /&gt;
this pre-existing error path does not mask Problem #263.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=823</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=823"/>
		<updated>2026-09-26T00:19:36Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Update details of Problem #112&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commits efa75af037a5e22b2b43675b435f472193a6672f and 921b5ca3e9ca3d1e06fe189b4775aae7dcc484bf&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
9/3/2026 - ICG: Actually this returns cases where a scientific food name is associated with more than one local food name. That is, There can be multiple words for the same latin name.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt; will be derived from&lt;br /&gt;
&amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt;, not &amp;lt;code&amp;gt;fl_sci_food_name&amp;lt;/code&amp;gt;.  Problem&lt;br /&gt;
#87 therefore controls description collisions for this conversion.&lt;br /&gt;
&lt;br /&gt;
The historical query above does not test the relationship stated in the&lt;br /&gt;
heading: it finds a scientific name shared by multiple local names.  Do not use&lt;br /&gt;
the historical count of 47 as a conversion assertion.  If this issue is&lt;br /&gt;
revisited, first replace the query with one that tests the intended&lt;br /&gt;
relationship against the refreshed source.&lt;br /&gt;
&lt;br /&gt;
== (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
9/4/2026 &lt;br /&gt;
Ian fixed mifupa, miti and mizizi in FOOD_BOUT by changing to UNRECORDED. DELETED FROM FOOD_PART_LOOKUP&lt;br /&gt;
Consolidated insects to &amp;#039;dudu&amp;#039;&lt;br /&gt;
changed all &amp;quot;NA&amp;quot; to &amp;quot;None&amp;quot;&lt;br /&gt;
Kept unrecorded&lt;br /&gt;
fixed spellings of utomvi and chipukizi&lt;br /&gt;
&lt;br /&gt;
I made all associated changes in FOOD_BOUT, choosing to use names rather than initials&lt;br /&gt;
&lt;br /&gt;
== (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Use &amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt; for&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt;.  Exclude lookup rows where&lt;br /&gt;
&amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt; is true before deriving descriptions.  Leave each&lt;br /&gt;
noncolliding generalized description unchanged.  When one generalized&lt;br /&gt;
description is shared by multiple eligible local names, derive each target&lt;br /&gt;
description as:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
fl_sci_food_name_gen || &amp;#039; -- &amp;#039; || fl_local_food_name&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This representation is stable and preserves both source values.  Do not use&lt;br /&gt;
mung&amp;#039;s row-order-dependent numeric suffixes.  In the 2026-09-13 snapshot, 39&lt;br /&gt;
generalized descriptions remained shared by 111 eligible local names after&lt;br /&gt;
applying &amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt;.  Rerun a corrected collision query as this&lt;br /&gt;
problem is encountered; the historical count of 42 predates the approved&lt;br /&gt;
exclusion rule.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Do not silently choose between a compound primary part and a conflicting&lt;br /&gt;
explicit second part.  If such a conflict remains after refreshed source&lt;br /&gt;
cleanup, exclude the exact source row and document all source columns in the&lt;br /&gt;
loader predicate.&lt;br /&gt;
&lt;br /&gt;
The dated 2026-09-13 source contained one conflict:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Date !! Focal !! Begin !! End !! Primary part !! Primary name !! Explicit second part !! Second name&lt;br /&gt;
|-&lt;br /&gt;
| 1985-10-26 || EV || 11:22 || 11:51 || MATUNDA; CHIPUKIZI || MBULA || MATUNDA; CHIPUKIZI || BISHURUSHURU&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Under the approved parser, the compound second token is&lt;br /&gt;
&amp;lt;code&amp;gt;CHIPUKIZI&amp;lt;/code&amp;gt;, while the explicit second value is the entire compound&lt;br /&gt;
&amp;lt;code&amp;gt;MATUNDA; CHIPUKIZI&amp;lt;/code&amp;gt;.  The older table above uses the stale spelling&lt;br /&gt;
&amp;lt;code&amp;gt;CHIPUKIZA&amp;lt;/code&amp;gt;.  Reproduce the complete current row exactly before&lt;br /&gt;
adding the exclusion; do not rely on this dated spelling or count.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Add &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt;, described as &amp;lt;code&amp;gt;No data recorded&amp;lt;/code&amp;gt;, to both&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;FOOD_PARTS&amp;lt;/code&amp;gt;.  Create a Seq 2 row when&lt;br /&gt;
either approved secondary component exists.  Use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; only for&lt;br /&gt;
the missing half of that pair:&lt;br /&gt;
&lt;br /&gt;
* part exists, name missing: use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; for FoodName;&lt;br /&gt;
* name exists, part missing: use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; for FoodPart;&lt;br /&gt;
* neither exists: create no Seq 2 row; and&lt;br /&gt;
* both exist: preserve both approved values.&lt;br /&gt;
&lt;br /&gt;
Normalize colon and semicolon delimiters and derive ordered compound tokens for&lt;br /&gt;
Seq 1 and Seq 2.  If a compound-derived second part conflicts with an explicit&lt;br /&gt;
second part, exactly exclude the row under Problem #89.  Do not silently apply&lt;br /&gt;
precedence.  The dated source had 12 secondary names without an explicit second&lt;br /&gt;
part; these are eligible for the approved &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; FoodPart rather&lt;br /&gt;
than exclusion.  Rerun both queries above with the approved clean normalization&lt;br /&gt;
and exclusion projection as this problem is encountered.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, add one deterministic spelling of each case-insensitive GROOM_SCANS extractor value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite gs_extracted_by to the exact spelling stored in clean.people.&lt;br /&gt;
&lt;br /&gt;
The ordinary support-table loader copies these rows to codes.people with active set to true before B-record groom scans are loaded. This preserves all 44,679 source rows and leaves the production foreign key and active-person rule intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of `U` in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
Approved by the investigator on 2026-09-11.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#136) Some GROOM_SCANS direction codes have trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 GROOM_SCANS records where GS_direction is &amp;#039;G &amp;#039; instead of &amp;#039;G&amp;#039;. The trailing Access padding prevents an exact match with the valid one-character direction code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in the clean schema, so query the easy schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;#039;&amp;quot;&amp;#039; || gs_direction || &amp;#039;&amp;quot;&amp;#039; AS untrimmed_direction,&lt;br /&gt;
       &amp;#039;&amp;quot;&amp;#039; || BTRIM(gs_direction) || &amp;#039;&amp;quot;&amp;#039; AS trimmed_direction,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM easy.groom_scans&lt;br /&gt;
 WHERE gs_direction IS DISTINCT FROM BTRIM(gs_direction)&lt;br /&gt;
 GROUP BY gs_direction,&lt;br /&gt;
          BTRIM(gs_direction)&lt;br /&gt;
 ORDER BY gs_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This reports 3 rows having &amp;quot;G &amp;quot;, which normalizes to &amp;quot;G&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Trim leading and trailing whitespace from GS_direction in the clean schema. This preserves all source rows and allows the direction codes to be mapped normally by the production loader.&lt;br /&gt;
&lt;br /&gt;
Resolved by commit 3ad4f5e1e2355895e5e465c4b8049197c8dce295&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=822</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=822"/>
		<updated>2026-09-26T00:18:33Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Update details of Problem #106&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commits efa75af037a5e22b2b43675b435f472193a6672f and 921b5ca3e9ca3d1e06fe189b4775aae7dcc484bf&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
9/3/2026 - ICG: Actually this returns cases where a scientific food name is associated with more than one local food name. That is, There can be multiple words for the same latin name.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt; will be derived from&lt;br /&gt;
&amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt;, not &amp;lt;code&amp;gt;fl_sci_food_name&amp;lt;/code&amp;gt;.  Problem&lt;br /&gt;
#87 therefore controls description collisions for this conversion.&lt;br /&gt;
&lt;br /&gt;
The historical query above does not test the relationship stated in the&lt;br /&gt;
heading: it finds a scientific name shared by multiple local names.  Do not use&lt;br /&gt;
the historical count of 47 as a conversion assertion.  If this issue is&lt;br /&gt;
revisited, first replace the query with one that tests the intended&lt;br /&gt;
relationship against the refreshed source.&lt;br /&gt;
&lt;br /&gt;
== (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
9/4/2026 &lt;br /&gt;
Ian fixed mifupa, miti and mizizi in FOOD_BOUT by changing to UNRECORDED. DELETED FROM FOOD_PART_LOOKUP&lt;br /&gt;
Consolidated insects to &amp;#039;dudu&amp;#039;&lt;br /&gt;
changed all &amp;quot;NA&amp;quot; to &amp;quot;None&amp;quot;&lt;br /&gt;
Kept unrecorded&lt;br /&gt;
fixed spellings of utomvi and chipukizi&lt;br /&gt;
&lt;br /&gt;
I made all associated changes in FOOD_BOUT, choosing to use names rather than initials&lt;br /&gt;
&lt;br /&gt;
== (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Use &amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt; for&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt;.  Exclude lookup rows where&lt;br /&gt;
&amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt; is true before deriving descriptions.  Leave each&lt;br /&gt;
noncolliding generalized description unchanged.  When one generalized&lt;br /&gt;
description is shared by multiple eligible local names, derive each target&lt;br /&gt;
description as:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
fl_sci_food_name_gen || &amp;#039; -- &amp;#039; || fl_local_food_name&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This representation is stable and preserves both source values.  Do not use&lt;br /&gt;
mung&amp;#039;s row-order-dependent numeric suffixes.  In the 2026-09-13 snapshot, 39&lt;br /&gt;
generalized descriptions remained shared by 111 eligible local names after&lt;br /&gt;
applying &amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt;.  Rerun a corrected collision query as this&lt;br /&gt;
problem is encountered; the historical count of 42 predates the approved&lt;br /&gt;
exclusion rule.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Do not silently choose between a compound primary part and a conflicting&lt;br /&gt;
explicit second part.  If such a conflict remains after refreshed source&lt;br /&gt;
cleanup, exclude the exact source row and document all source columns in the&lt;br /&gt;
loader predicate.&lt;br /&gt;
&lt;br /&gt;
The dated 2026-09-13 source contained one conflict:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Date !! Focal !! Begin !! End !! Primary part !! Primary name !! Explicit second part !! Second name&lt;br /&gt;
|-&lt;br /&gt;
| 1985-10-26 || EV || 11:22 || 11:51 || MATUNDA; CHIPUKIZI || MBULA || MATUNDA; CHIPUKIZI || BISHURUSHURU&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Under the approved parser, the compound second token is&lt;br /&gt;
&amp;lt;code&amp;gt;CHIPUKIZI&amp;lt;/code&amp;gt;, while the explicit second value is the entire compound&lt;br /&gt;
&amp;lt;code&amp;gt;MATUNDA; CHIPUKIZI&amp;lt;/code&amp;gt;.  The older table above uses the stale spelling&lt;br /&gt;
&amp;lt;code&amp;gt;CHIPUKIZA&amp;lt;/code&amp;gt;.  Reproduce the complete current row exactly before&lt;br /&gt;
adding the exclusion; do not rely on this dated spelling or count.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Add &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt;, described as &amp;lt;code&amp;gt;No data recorded&amp;lt;/code&amp;gt;, to both&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;FOOD_PARTS&amp;lt;/code&amp;gt;.  Create a Seq 2 row when&lt;br /&gt;
either approved secondary component exists.  Use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; only for&lt;br /&gt;
the missing half of that pair:&lt;br /&gt;
&lt;br /&gt;
* part exists, name missing: use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; for FoodName;&lt;br /&gt;
* name exists, part missing: use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; for FoodPart;&lt;br /&gt;
* neither exists: create no Seq 2 row; and&lt;br /&gt;
* both exist: preserve both approved values.&lt;br /&gt;
&lt;br /&gt;
Normalize colon and semicolon delimiters and derive ordered compound tokens for&lt;br /&gt;
Seq 1 and Seq 2.  If a compound-derived second part conflicts with an explicit&lt;br /&gt;
second part, exactly exclude the row under Problem #89.  Do not silently apply&lt;br /&gt;
precedence.  The dated source had 12 secondary names without an explicit second&lt;br /&gt;
part; these are eligible for the approved &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; FoodPart rather&lt;br /&gt;
than exclusion.  Rerun both queries above with the approved clean normalization&lt;br /&gt;
and exclusion projection as this problem is encountered.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, add one deterministic spelling of each case-insensitive GROOM_SCANS extractor value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite gs_extracted_by to the exact spelling stored in clean.people.&lt;br /&gt;
&lt;br /&gt;
The ordinary support-table loader copies these rows to codes.people with active set to true before B-record groom scans are loaded. This preserves all 44,679 source rows and leaves the production foreign key and active-person rule intact.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#136) Some GROOM_SCANS direction codes have trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 GROOM_SCANS records where GS_direction is &amp;#039;G &amp;#039; instead of &amp;#039;G&amp;#039;. The trailing Access padding prevents an exact match with the valid one-character direction code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in the clean schema, so query the easy schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;#039;&amp;quot;&amp;#039; || gs_direction || &amp;#039;&amp;quot;&amp;#039; AS untrimmed_direction,&lt;br /&gt;
       &amp;#039;&amp;quot;&amp;#039; || BTRIM(gs_direction) || &amp;#039;&amp;quot;&amp;#039; AS trimmed_direction,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM easy.groom_scans&lt;br /&gt;
 WHERE gs_direction IS DISTINCT FROM BTRIM(gs_direction)&lt;br /&gt;
 GROUP BY gs_direction,&lt;br /&gt;
          BTRIM(gs_direction)&lt;br /&gt;
 ORDER BY gs_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This reports 3 rows having &amp;quot;G &amp;quot;, which normalizes to &amp;quot;G&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Trim leading and trailing whitespace from GS_direction in the clean schema. This preserves all source rows and allows the direction codes to be mapped normally by the production loader.&lt;br /&gt;
&lt;br /&gt;
Resolved by commit 3ad4f5e1e2355895e5e465c4b8049197c8dce295&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=821</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=821"/>
		<updated>2026-09-26T00:15:49Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Update details of Problem #93&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commits efa75af037a5e22b2b43675b435f472193a6672f and 921b5ca3e9ca3d1e06fe189b4775aae7dcc484bf&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
9/3/2026 - ICG: Actually this returns cases where a scientific food name is associated with more than one local food name. That is, There can be multiple words for the same latin name.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt; will be derived from&lt;br /&gt;
&amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt;, not &amp;lt;code&amp;gt;fl_sci_food_name&amp;lt;/code&amp;gt;.  Problem&lt;br /&gt;
#87 therefore controls description collisions for this conversion.&lt;br /&gt;
&lt;br /&gt;
The historical query above does not test the relationship stated in the&lt;br /&gt;
heading: it finds a scientific name shared by multiple local names.  Do not use&lt;br /&gt;
the historical count of 47 as a conversion assertion.  If this issue is&lt;br /&gt;
revisited, first replace the query with one that tests the intended&lt;br /&gt;
relationship against the refreshed source.&lt;br /&gt;
&lt;br /&gt;
== (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
9/4/2026 &lt;br /&gt;
Ian fixed mifupa, miti and mizizi in FOOD_BOUT by changing to UNRECORDED. DELETED FROM FOOD_PART_LOOKUP&lt;br /&gt;
Consolidated insects to &amp;#039;dudu&amp;#039;&lt;br /&gt;
changed all &amp;quot;NA&amp;quot; to &amp;quot;None&amp;quot;&lt;br /&gt;
Kept unrecorded&lt;br /&gt;
fixed spellings of utomvi and chipukizi&lt;br /&gt;
&lt;br /&gt;
I made all associated changes in FOOD_BOUT, choosing to use names rather than initials&lt;br /&gt;
&lt;br /&gt;
== (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Use &amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt; for&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt;.  Exclude lookup rows where&lt;br /&gt;
&amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt; is true before deriving descriptions.  Leave each&lt;br /&gt;
noncolliding generalized description unchanged.  When one generalized&lt;br /&gt;
description is shared by multiple eligible local names, derive each target&lt;br /&gt;
description as:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
fl_sci_food_name_gen || &amp;#039; -- &amp;#039; || fl_local_food_name&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This representation is stable and preserves both source values.  Do not use&lt;br /&gt;
mung&amp;#039;s row-order-dependent numeric suffixes.  In the 2026-09-13 snapshot, 39&lt;br /&gt;
generalized descriptions remained shared by 111 eligible local names after&lt;br /&gt;
applying &amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt;.  Rerun a corrected collision query as this&lt;br /&gt;
problem is encountered; the historical count of 42 predates the approved&lt;br /&gt;
exclusion rule.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Do not silently choose between a compound primary part and a conflicting&lt;br /&gt;
explicit second part.  If such a conflict remains after refreshed source&lt;br /&gt;
cleanup, exclude the exact source row and document all source columns in the&lt;br /&gt;
loader predicate.&lt;br /&gt;
&lt;br /&gt;
The dated 2026-09-13 source contained one conflict:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Date !! Focal !! Begin !! End !! Primary part !! Primary name !! Explicit second part !! Second name&lt;br /&gt;
|-&lt;br /&gt;
| 1985-10-26 || EV || 11:22 || 11:51 || MATUNDA; CHIPUKIZI || MBULA || MATUNDA; CHIPUKIZI || BISHURUSHURU&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Under the approved parser, the compound second token is&lt;br /&gt;
&amp;lt;code&amp;gt;CHIPUKIZI&amp;lt;/code&amp;gt;, while the explicit second value is the entire compound&lt;br /&gt;
&amp;lt;code&amp;gt;MATUNDA; CHIPUKIZI&amp;lt;/code&amp;gt;.  The older table above uses the stale spelling&lt;br /&gt;
&amp;lt;code&amp;gt;CHIPUKIZA&amp;lt;/code&amp;gt;.  Reproduce the complete current row exactly before&lt;br /&gt;
adding the exclusion; do not rely on this dated spelling or count.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Add &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt;, described as &amp;lt;code&amp;gt;No data recorded&amp;lt;/code&amp;gt;, to both&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;FOOD_PARTS&amp;lt;/code&amp;gt;.  Create a Seq 2 row when&lt;br /&gt;
either approved secondary component exists.  Use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; only for&lt;br /&gt;
the missing half of that pair:&lt;br /&gt;
&lt;br /&gt;
* part exists, name missing: use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; for FoodName;&lt;br /&gt;
* name exists, part missing: use &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; for FoodPart;&lt;br /&gt;
* neither exists: create no Seq 2 row; and&lt;br /&gt;
* both exist: preserve both approved values.&lt;br /&gt;
&lt;br /&gt;
Normalize colon and semicolon delimiters and derive ordered compound tokens for&lt;br /&gt;
Seq 1 and Seq 2.  If a compound-derived second part conflicts with an explicit&lt;br /&gt;
second part, exactly exclude the row under Problem #89.  Do not silently apply&lt;br /&gt;
precedence.  The dated source had 12 secondary names without an explicit second&lt;br /&gt;
part; these are eligible for the approved &amp;lt;code&amp;gt;NODATA&amp;lt;/code&amp;gt; FoodPart rather&lt;br /&gt;
than exclusion.  Rerun both queries above with the approved clean normalization&lt;br /&gt;
and exclusion projection as this problem is encountered.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#136) Some GROOM_SCANS direction codes have trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 GROOM_SCANS records where GS_direction is &amp;#039;G &amp;#039; instead of &amp;#039;G&amp;#039;. The trailing Access padding prevents an exact match with the valid one-character direction code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in the clean schema, so query the easy schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;#039;&amp;quot;&amp;#039; || gs_direction || &amp;#039;&amp;quot;&amp;#039; AS untrimmed_direction,&lt;br /&gt;
       &amp;#039;&amp;quot;&amp;#039; || BTRIM(gs_direction) || &amp;#039;&amp;quot;&amp;#039; AS trimmed_direction,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM easy.groom_scans&lt;br /&gt;
 WHERE gs_direction IS DISTINCT FROM BTRIM(gs_direction)&lt;br /&gt;
 GROUP BY gs_direction,&lt;br /&gt;
          BTRIM(gs_direction)&lt;br /&gt;
 ORDER BY gs_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This reports 3 rows having &amp;quot;G &amp;quot;, which normalizes to &amp;quot;G&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Trim leading and trailing whitespace from GS_direction in the clean schema. This preserves all source rows and allows the direction codes to be mapped normally by the production loader.&lt;br /&gt;
&lt;br /&gt;
Resolved by commit 3ad4f5e1e2355895e5e465c4b8049197c8dce295&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=820</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=820"/>
		<updated>2026-09-26T00:13:44Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Update details of Problem #89&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commits efa75af037a5e22b2b43675b435f472193a6672f and 921b5ca3e9ca3d1e06fe189b4775aae7dcc484bf&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
9/3/2026 - ICG: Actually this returns cases where a scientific food name is associated with more than one local food name. That is, There can be multiple words for the same latin name.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt; will be derived from&lt;br /&gt;
&amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt;, not &amp;lt;code&amp;gt;fl_sci_food_name&amp;lt;/code&amp;gt;.  Problem&lt;br /&gt;
#87 therefore controls description collisions for this conversion.&lt;br /&gt;
&lt;br /&gt;
The historical query above does not test the relationship stated in the&lt;br /&gt;
heading: it finds a scientific name shared by multiple local names.  Do not use&lt;br /&gt;
the historical count of 47 as a conversion assertion.  If this issue is&lt;br /&gt;
revisited, first replace the query with one that tests the intended&lt;br /&gt;
relationship against the refreshed source.&lt;br /&gt;
&lt;br /&gt;
== (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
9/4/2026 &lt;br /&gt;
Ian fixed mifupa, miti and mizizi in FOOD_BOUT by changing to UNRECORDED. DELETED FROM FOOD_PART_LOOKUP&lt;br /&gt;
Consolidated insects to &amp;#039;dudu&amp;#039;&lt;br /&gt;
changed all &amp;quot;NA&amp;quot; to &amp;quot;None&amp;quot;&lt;br /&gt;
Kept unrecorded&lt;br /&gt;
fixed spellings of utomvi and chipukizi&lt;br /&gt;
&lt;br /&gt;
I made all associated changes in FOOD_BOUT, choosing to use names rather than initials&lt;br /&gt;
&lt;br /&gt;
== (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Use &amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt; for&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt;.  Exclude lookup rows where&lt;br /&gt;
&amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt; is true before deriving descriptions.  Leave each&lt;br /&gt;
noncolliding generalized description unchanged.  When one generalized&lt;br /&gt;
description is shared by multiple eligible local names, derive each target&lt;br /&gt;
description as:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
fl_sci_food_name_gen || &amp;#039; -- &amp;#039; || fl_local_food_name&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This representation is stable and preserves both source values.  Do not use&lt;br /&gt;
mung&amp;#039;s row-order-dependent numeric suffixes.  In the 2026-09-13 snapshot, 39&lt;br /&gt;
generalized descriptions remained shared by 111 eligible local names after&lt;br /&gt;
applying &amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt;.  Rerun a corrected collision query as this&lt;br /&gt;
problem is encountered; the historical count of 42 predates the approved&lt;br /&gt;
exclusion rule.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Do not silently choose between a compound primary part and a conflicting&lt;br /&gt;
explicit second part.  If such a conflict remains after refreshed source&lt;br /&gt;
cleanup, exclude the exact source row and document all source columns in the&lt;br /&gt;
loader predicate.&lt;br /&gt;
&lt;br /&gt;
The dated 2026-09-13 source contained one conflict:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Date !! Focal !! Begin !! End !! Primary part !! Primary name !! Explicit second part !! Second name&lt;br /&gt;
|-&lt;br /&gt;
| 1985-10-26 || EV || 11:22 || 11:51 || MATUNDA; CHIPUKIZI || MBULA || MATUNDA; CHIPUKIZI || BISHURUSHURU&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Under the approved parser, the compound second token is&lt;br /&gt;
&amp;lt;code&amp;gt;CHIPUKIZI&amp;lt;/code&amp;gt;, while the explicit second value is the entire compound&lt;br /&gt;
&amp;lt;code&amp;gt;MATUNDA; CHIPUKIZI&amp;lt;/code&amp;gt;.  The older table above uses the stale spelling&lt;br /&gt;
&amp;lt;code&amp;gt;CHIPUKIZA&amp;lt;/code&amp;gt;.  Reproduce the complete current row exactly before&lt;br /&gt;
adding the exclusion; do not rely on this dated spelling or count.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#136) Some GROOM_SCANS direction codes have trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 GROOM_SCANS records where GS_direction is &amp;#039;G &amp;#039; instead of &amp;#039;G&amp;#039;. The trailing Access padding prevents an exact match with the valid one-character direction code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in the clean schema, so query the easy schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;#039;&amp;quot;&amp;#039; || gs_direction || &amp;#039;&amp;quot;&amp;#039; AS untrimmed_direction,&lt;br /&gt;
       &amp;#039;&amp;quot;&amp;#039; || BTRIM(gs_direction) || &amp;#039;&amp;quot;&amp;#039; AS trimmed_direction,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM easy.groom_scans&lt;br /&gt;
 WHERE gs_direction IS DISTINCT FROM BTRIM(gs_direction)&lt;br /&gt;
 GROUP BY gs_direction,&lt;br /&gt;
          BTRIM(gs_direction)&lt;br /&gt;
 ORDER BY gs_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This reports 3 rows having &amp;quot;G &amp;quot;, which normalizes to &amp;quot;G&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Trim leading and trailing whitespace from GS_direction in the clean schema. This preserves all source rows and allows the direction codes to be mapped normally by the production loader.&lt;br /&gt;
&lt;br /&gt;
Resolved by commit 3ad4f5e1e2355895e5e465c4b8049197c8dce295&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=819</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=819"/>
		<updated>2026-09-26T00:12:30Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Update details of Problem #87&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commits efa75af037a5e22b2b43675b435f472193a6672f and 921b5ca3e9ca3d1e06fe189b4775aae7dcc484bf&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
9/3/2026 - ICG: Actually this returns cases where a scientific food name is associated with more than one local food name. That is, There can be multiple words for the same latin name.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt; will be derived from&lt;br /&gt;
&amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt;, not &amp;lt;code&amp;gt;fl_sci_food_name&amp;lt;/code&amp;gt;.  Problem&lt;br /&gt;
#87 therefore controls description collisions for this conversion.&lt;br /&gt;
&lt;br /&gt;
The historical query above does not test the relationship stated in the&lt;br /&gt;
heading: it finds a scientific name shared by multiple local names.  Do not use&lt;br /&gt;
the historical count of 47 as a conversion assertion.  If this issue is&lt;br /&gt;
revisited, first replace the query with one that tests the intended&lt;br /&gt;
relationship against the refreshed source.&lt;br /&gt;
&lt;br /&gt;
== (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
9/4/2026 &lt;br /&gt;
Ian fixed mifupa, miti and mizizi in FOOD_BOUT by changing to UNRECORDED. DELETED FROM FOOD_PART_LOOKUP&lt;br /&gt;
Consolidated insects to &amp;#039;dudu&amp;#039;&lt;br /&gt;
changed all &amp;quot;NA&amp;quot; to &amp;quot;None&amp;quot;&lt;br /&gt;
Kept unrecorded&lt;br /&gt;
fixed spellings of utomvi and chipukizi&lt;br /&gt;
&lt;br /&gt;
I made all associated changes in FOOD_BOUT, choosing to use names rather than initials&lt;br /&gt;
&lt;br /&gt;
== (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
Use &amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt; for&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt;.  Exclude lookup rows where&lt;br /&gt;
&amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt; is true before deriving descriptions.  Leave each&lt;br /&gt;
noncolliding generalized description unchanged.  When one generalized&lt;br /&gt;
description is shared by multiple eligible local names, derive each target&lt;br /&gt;
description as:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
fl_sci_food_name_gen || &amp;#039; -- &amp;#039; || fl_local_food_name&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This representation is stable and preserves both source values.  Do not use&lt;br /&gt;
mung&amp;#039;s row-order-dependent numeric suffixes.  In the 2026-09-13 snapshot, 39&lt;br /&gt;
generalized descriptions remained shared by 111 eligible local names after&lt;br /&gt;
applying &amp;lt;code&amp;gt;FL_exclude&amp;lt;/code&amp;gt;.  Rerun a corrected collision query as this&lt;br /&gt;
problem is encountered; the historical count of 42 predates the approved&lt;br /&gt;
exclusion rule.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#136) Some GROOM_SCANS direction codes have trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 GROOM_SCANS records where GS_direction is &amp;#039;G &amp;#039; instead of &amp;#039;G&amp;#039;. The trailing Access padding prevents an exact match with the valid one-character direction code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in the clean schema, so query the easy schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;#039;&amp;quot;&amp;#039; || gs_direction || &amp;#039;&amp;quot;&amp;#039; AS untrimmed_direction,&lt;br /&gt;
       &amp;#039;&amp;quot;&amp;#039; || BTRIM(gs_direction) || &amp;#039;&amp;quot;&amp;#039; AS trimmed_direction,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM easy.groom_scans&lt;br /&gt;
 WHERE gs_direction IS DISTINCT FROM BTRIM(gs_direction)&lt;br /&gt;
 GROUP BY gs_direction,&lt;br /&gt;
          BTRIM(gs_direction)&lt;br /&gt;
 ORDER BY gs_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This reports 3 rows having &amp;quot;G &amp;quot;, which normalizes to &amp;quot;G&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Trim leading and trailing whitespace from GS_direction in the clean schema. This preserves all source rows and allows the direction codes to be mapped normally by the production loader.&lt;br /&gt;
&lt;br /&gt;
Resolved by commit 3ad4f5e1e2355895e5e465c4b8049197c8dce295&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=818</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=818"/>
		<updated>2026-09-26T00:10:43Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Update details of Problem #85&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commits efa75af037a5e22b2b43675b435f472193a6672f and 921b5ca3e9ca3d1e06fe189b4775aae7dcc484bf&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
9/3/2026 - ICG: Actually this returns cases where a scientific food name is associated with more than one local food name. That is, There can be multiple words for the same latin name.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
==== 2026-09-13 conversion decision ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;FOOD_NAMES.Description&amp;lt;/code&amp;gt; will be derived from&lt;br /&gt;
&amp;lt;code&amp;gt;fl_sci_food_name_gen&amp;lt;/code&amp;gt;, not &amp;lt;code&amp;gt;fl_sci_food_name&amp;lt;/code&amp;gt;.  Problem&lt;br /&gt;
#87 therefore controls description collisions for this conversion.&lt;br /&gt;
&lt;br /&gt;
The historical query above does not test the relationship stated in the&lt;br /&gt;
heading: it finds a scientific name shared by multiple local names.  Do not use&lt;br /&gt;
the historical count of 47 as a conversion assertion.  If this issue is&lt;br /&gt;
revisited, first replace the query with one that tests the intended&lt;br /&gt;
relationship against the refreshed source.&lt;br /&gt;
&lt;br /&gt;
== (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
9/4/2026 &lt;br /&gt;
Ian fixed mifupa, miti and mizizi in FOOD_BOUT by changing to UNRECORDED. DELETED FROM FOOD_PART_LOOKUP&lt;br /&gt;
Consolidated insects to &amp;#039;dudu&amp;#039;&lt;br /&gt;
changed all &amp;quot;NA&amp;quot; to &amp;quot;None&amp;quot;&lt;br /&gt;
Kept unrecorded&lt;br /&gt;
fixed spellings of utomvi and chipukizi&lt;br /&gt;
&lt;br /&gt;
I made all associated changes in FOOD_BOUT, choosing to use names rather than initials&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#136) Some GROOM_SCANS direction codes have trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 GROOM_SCANS records where GS_direction is &amp;#039;G &amp;#039; instead of &amp;#039;G&amp;#039;. The trailing Access padding prevents an exact match with the valid one-character direction code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in the clean schema, so query the easy schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;#039;&amp;quot;&amp;#039; || gs_direction || &amp;#039;&amp;quot;&amp;#039; AS untrimmed_direction,&lt;br /&gt;
       &amp;#039;&amp;quot;&amp;#039; || BTRIM(gs_direction) || &amp;#039;&amp;quot;&amp;#039; AS trimmed_direction,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM easy.groom_scans&lt;br /&gt;
 WHERE gs_direction IS DISTINCT FROM BTRIM(gs_direction)&lt;br /&gt;
 GROUP BY gs_direction,&lt;br /&gt;
          BTRIM(gs_direction)&lt;br /&gt;
 ORDER BY gs_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This reports 3 rows having &amp;quot;G &amp;quot;, which normalizes to &amp;quot;G&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Trim leading and trailing whitespace from GS_direction in the clean schema. This preserves all source rows and allows the direction codes to be mapped normally by the production loader.&lt;br /&gt;
&lt;br /&gt;
Resolved by commit 3ad4f5e1e2355895e5e465c4b8049197c8dce295&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=817</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=817"/>
		<updated>2026-09-11T03:06:35Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Problem #136 Some GROOM_SCANS direction codes have trailing spaces&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commits efa75af037a5e22b2b43675b435f472193a6672f and 921b5ca3e9ca3d1e06fe189b4775aae7dcc484bf&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
9/3/2026 - ICG: Actually this returns cases where a scientific food name is associated with more than one local food name. That is, There can be multiple words for the same latin name.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
9/4/2026 &lt;br /&gt;
Ian fixed mifupa, miti and mizizi in FOOD_BOUT by changing to UNRECORDED. DELETED FROM FOOD_PART_LOOKUP&lt;br /&gt;
Consolidated insects to &amp;#039;dudu&amp;#039;&lt;br /&gt;
changed all &amp;quot;NA&amp;quot; to &amp;quot;None&amp;quot;&lt;br /&gt;
Kept unrecorded&lt;br /&gt;
fixed spellings of utomvi and chipukizi&lt;br /&gt;
&lt;br /&gt;
I made all associated changes in FOOD_BOUT, choosing to use names rather than initials&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#136) Some GROOM_SCANS direction codes have trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 GROOM_SCANS records where GS_direction is &amp;#039;G &amp;#039; instead of &amp;#039;G&amp;#039;. The trailing Access padding prevents an exact match with the valid one-character direction code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in the clean schema, so query the easy schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;#039;&amp;quot;&amp;#039; || gs_direction || &amp;#039;&amp;quot;&amp;#039; AS untrimmed_direction,&lt;br /&gt;
       &amp;#039;&amp;quot;&amp;#039; || BTRIM(gs_direction) || &amp;#039;&amp;quot;&amp;#039; AS trimmed_direction,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM easy.groom_scans&lt;br /&gt;
 WHERE gs_direction IS DISTINCT FROM BTRIM(gs_direction)&lt;br /&gt;
 GROUP BY gs_direction,&lt;br /&gt;
          BTRIM(gs_direction)&lt;br /&gt;
 ORDER BY gs_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This reports 3 rows having &amp;quot;G &amp;quot;, which normalizes to &amp;quot;G&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Trim leading and trailing whitespace from GS_direction in the clean schema. This preserves all source rows and allows the direction codes to be mapped normally by the production loader.&lt;br /&gt;
&lt;br /&gt;
Resolved by commit 3ad4f5e1e2355895e5e465c4b8049197c8dce295&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=816</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=816"/>
		<updated>2026-09-10T03:34:13Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: fully addressing Problem #74 required another commit&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/24/2016&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 8/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commits efa75af037a5e22b2b43675b435f472193a6672f and 921b5ca3e9ca3d1e06fe189b4775aae7dcc484bf&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
9/3/2026 - ICG: Actually this returns cases where a scientific food name is associated with more than one local food name. That is, There can be multiple words for the same latin name.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
9/4/2026 &lt;br /&gt;
Ian fixed mifupa, miti and mizizi in FOOD_BOUT by changing to UNRECORDED. DELETED FROM FOOD_PART_LOOKUP&lt;br /&gt;
Consolidated insects to &amp;#039;dudu&amp;#039;&lt;br /&gt;
changed all &amp;quot;NA&amp;quot; to &amp;quot;None&amp;quot;&lt;br /&gt;
Kept unrecorded&lt;br /&gt;
fixed spellings of utomvi and chipukizi&lt;br /&gt;
&lt;br /&gt;
I made all associated changes in FOOD_BOUT, choosing to use names rather than initials&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=810</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=810"/>
		<updated>2026-08-22T00:14:47Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: clarify issue #130&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Partially resolved; unresolved for twelve no-focal/date keys. The current loader exclusion is broader than necessary.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt is related through EVENTS to one supporting WATCHES row. WATCHES records one CommID, and only one watch of a given Type may exist for a particular Date and AnimID. PANTGRUNTS also records its own CommID. This duplication is intentional: PANTGRUNTS.CommID preserves the community recorded with the pantgrunt in Access, while WATCHES.CommID records the community of the supporting follow or watch period. The two values are permitted to disagree, although the production database reports the disagreement as a warning.&lt;br /&gt;
&lt;br /&gt;
The local conversion data examined on 2026-08-21 contain 15 normalized focal/date keys with both KK and MT pantgrunt communities. These keys cover 104 pantgrunt rows. They fall into two materially different categories:&lt;br /&gt;
&lt;br /&gt;
* Three named-focal keys, covering 48 pantgrunt rows, already have a corresponding follow. The follow supplies an unambiguous B-record WATCHES row whose community is KK. These rows can be converted without choosing or discarding a pantgrunt community: reuse the existing watch and preserve each source value in PANTGRUNTS.CommID. Pantgrunts recorded as MT will appropriately produce community-mismatch warnings.&lt;br /&gt;
* Twelve keys, covering 56 pantgrunt rows, have no focal animal ID and no corresponding follow. The conversion normalizes the absent focal to NONE for watch resolution. Each date contains both KK and MT pantgrunts, but the conversion can create only one Other watch for the combination of NONE and that date. There is no source-supported way to decide whether that WATCHES row should contain KK or MT.&lt;br /&gt;
&lt;br /&gt;
The community conflict therefore does not itself prevent a PANTGRUNTS row from storing its source community. The unresolved problem is selecting the canonical WATCHES.CommID when a supporting watch must be created and there is no follow or focal context from which to determine it.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query shows every conflicting key, any community supplied by a corresponding follow, the pantgrunt communities, and the number of affected pantgrunt rows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH conflict_keys AS (&lt;br /&gt;
  SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
         pg_date&lt;br /&gt;
    FROM clean.pantgrunt_event&lt;br /&gt;
   GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
            pg_date&lt;br /&gt;
  HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
)&lt;br /&gt;
SELECT conflict_keys.watch_animid,&lt;br /&gt;
       conflict_keys.pg_date,&lt;br /&gt;
       follow.fol_cl_community_id AS follow_community,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pantgrunt_event.pg_cl_community_id)&lt;br /&gt;
                 ORDER BY BTRIM(pantgrunt_event.pg_cl_community_id)) AS pantgrunt_communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM conflict_keys&lt;br /&gt;
  JOIN clean.pantgrunt_event&lt;br /&gt;
    ON COALESCE(NULLIF(BTRIM(pantgrunt_event.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) = conflict_keys.watch_animid&lt;br /&gt;
       AND pantgrunt_event.pg_date = conflict_keys.pg_date&lt;br /&gt;
  LEFT JOIN clean.follow&lt;br /&gt;
    ON BTRIM(follow.fol_b_animid) = conflict_keys.watch_animid&lt;br /&gt;
       AND follow.fol_date = conflict_keys.pg_date&lt;br /&gt;
 GROUP BY conflict_keys.watch_animid,&lt;br /&gt;
          conflict_keys.pg_date,&lt;br /&gt;
          follow.fol_cl_community_id&lt;br /&gt;
 ORDER BY conflict_keys.pg_date,&lt;br /&gt;
          conflict_keys.watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Examples with named focals ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Focal&lt;br /&gt;
! Date&lt;br /&gt;
! Follow community&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| FD&lt;br /&gt;
| 1996-09-08&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 11&lt;br /&gt;
|-&lt;br /&gt;
| ZS&lt;br /&gt;
| 2013-11-22&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 6&lt;br /&gt;
|-&lt;br /&gt;
| DL&lt;br /&gt;
| 2013-11-29&lt;br /&gt;
| KK&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 31&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the FD records can all use the existing B-record watch for FD on 1996-09-08. That watch remains associated with KK. Each pantgrunt retains either KK or MT in PANTGRUNTS.CommID. The MT records disagree with the watch community, but this is a documented and intentionally preserved source inconsistency rather than a conversion failure.&lt;br /&gt;
&lt;br /&gt;
=== Examples without a focal ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Watch AnimID used by conversion&lt;br /&gt;
! Date&lt;br /&gt;
! Pantgrunt communities&lt;br /&gt;
! Pantgrunt rows&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-01-11&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 8&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-09-10&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 16&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1986-10-30&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1987-09-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 7&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1992-11-21&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1993-12-25&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-01-19&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-05-26&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-07-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 2&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-01&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-08-08&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 4&lt;br /&gt;
|-&lt;br /&gt;
| NONE&lt;br /&gt;
| 1994-10-07&lt;br /&gt;
| KK and MT&lt;br /&gt;
| 3&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
For example, the eight records on 1986-01-11 require a supporting Other watch with AnimID NONE. The WATCHES uniqueness rule permits only one such Other watch on that date. Assigning KK would be unsupported for the MT observations, while assigning MT would be unsupported for the KK observations. Creating two Other watches distinguished only by community is not permitted by the current WATCHES key.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Refine the conversion so that the three named-focal keys reuse their existing B-record watches and all 48 pantgrunt rows are loaded. Preserve the source community in PANTGRUNTS.CommID and accept the documented warning when it differs from the B-record watch community.&lt;br /&gt;
&lt;br /&gt;
Continue to exclude only the 56 pantgrunt rows belonging to the twelve no-focal/date keys. Report those twelve keys in the sanity check. The investigators must determine whether each date represents one observation period with incorrect community coding, two distinct observation periods that the current WATCHES key cannot represent, or records with incorrect dates or focal values.&lt;br /&gt;
&lt;br /&gt;
Until that refinement is implemented, the current loader conservatively excludes all 15 conflicting keys and therefore excludes 48 otherwise convertible named-focal rows in addition to the 56 genuinely ambiguous no-focal rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=809</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=809"/>
		<updated>2026-08-21T22:34:24Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Add pantgrunt problems #118 - #135&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;br /&gt;
&lt;br /&gt;
== (#118) PANTGRUNT_EVENT rows have no recorded time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some PANTGRUNT_EVENT rows have NULL in pg_time. EVENTS.Start and EVENTS.Stop are not nullable, but the absence of a recorded pantgrunt time is meaningful and must not cause the entire row to be discarded. In the local conversion data examined on 2026-08-21, 684 rows had no recorded time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL pg_time to the sdb_no_time value, midnight, and use that value for both EVENTS.Start and EVENTS.Stop. Extend the EVENTS constraints so pantgrunt events, like aggression events, may use sdb_no_time. Retain NULL in clean.pantgrunt_event so the clean schema continues to distinguish missing source data from an actual recorded time.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#119) PANTGRUNT_EVENT times may contain seconds ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS times are stored to the minute and reject nonzero seconds. The Access pg_time value is initially represented as a timestamp, and merely converting it to a PostgreSQL time value does not remove seconds. The local clean data examined on 2026-08-21 contained no remaining nonzero seconds, but the pantgrunt workflow previously had no explicit normalization guaranteeing that result.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_date,&lt;br /&gt;
       pg_fol_b_foc_id,&lt;br /&gt;
       pg_time,&lt;br /&gt;
       extract(second FROM pg_time) AS seconds&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE extract(second FROM pg_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the tidy schema, truncate pg_time to the minute before tidy.sql converts the column to TIME. Do not round the value. The pantgrunt sanity check rejects the conversion if a value containing seconds nevertheless reaches the clean schema.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#120) PANTGRUNT_EVENT times fall outside the permitted event window ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for sdb_no_time, production event times must be between 04:00 and 20:00, inclusive. In the local conversion data examined on 2026-08-21, 35 non-NULL pantgrunt times were outside this interval. Their correct values cannot be inferred during conversion.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_time IS NOT NULL&lt;br /&gt;
       AND (pg_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR pg_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion and report their count in the pantgrunt sanity check. Retain the rows unchanged in clean.pantgrunt_event. The project investigators must review and correct the source times in Access; the conversion must not clamp them or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#121) PANTGRUNT_EVENT rows name the same actor and recipient ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production pantgrunt is a dyadic event involving two different participants. Some source rows identify the same animal as both pg_b_actor_id and pg_b_recipient_id. In the local conversion data examined on 2026-08-21, 11 rows had the same trimmed actor and recipient value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE BTRIM(pg_b_actor_id) = BTRIM(pg_b_recipient_id)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the pantgrunt sanity check. The project investigators must determine the correct actor or recipient and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#122) PANTGRUNT_EVENT rows have no data source ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.Source is required and must reference PG_SOURCES, but some source pantgrunt rows have a NULL, empty, or whitespace-only pg_data_source. In the local conversion data examined on 2026-08-21, 109 rows had no data source.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_data_source), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, normalize source codes by trimming and converting them to uppercase. Convert a missing source to NODATA. Populate PG_SOURCES from the distinct normalized source values and give NODATA the description &amp;quot;No pantgrunt data source was recorded&amp;quot;. Load NODATA into PANTGRUNTS.Source for affected rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#123) PANTGRUNT_EVENT rows have no extractor ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.EnteredBy is required and must reference an active PEOPLE row, but some source rows have a NULL, empty, or whitespace-only pg_extracted_by value. In the local conversion data examined on 2026-08-21, 2,223 rows had no recorded extractor.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, change a NULL, empty, or whitespace-only pg_extracted_by value to UNK. The established UNK row is copied from clean.people to codes.people as an active person and is loaded into PANTGRUNTS.EnteredBy for affected eligible rows.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#124) PANTGRUNT_EVENT extractor values are missing from PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Many nonempty pg_extracted_by values do not exactly match a clean.people person value. The local conversion data examined on 2026-08-21 had 26,334 rows without an exact PEOPLE match when missing values were included, and 24,111 nonmissing rows without a match after trimming. Some extractor names also occur with case variants, such as Anne and ANNE, while codes.people requires case-insensitive uniqueness.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT BTRIM(pg.pg_extracted_by) AS extracted_by,&lt;br /&gt;
       COUNT(*) AS row_count,&lt;br /&gt;
       ARRAY_AGG(DISTINCT pg.pg_extracted_by ORDER BY pg.pg_extracted_by) AS source_spellings&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.people AS p&lt;br /&gt;
          WHERE p.person = BTRIM(pg.pg_extracted_by))&lt;br /&gt;
 GROUP BY BTRIM(pg.pg_extracted_by)&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, trim nonempty pantgrunt extractor values and add one deterministic spelling of each case-insensitive value to clean.people when no case-insensitive PEOPLE match already exists. Prefer an existing PEOPLE spelling, and otherwise prefer a mixed-case source spelling over an all-uppercase or all-lowercase spelling. Rewrite pg_extracted_by to the exact spelling stored in clean.people. The ordinary support-table loader then copies these rows to codes.people with active set to true, which is required because new PANTGRUNTS rows may reference only active people.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#125) PANTGRUNT_EVENT two-sided flags use multiple encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_two_sided_flag determines whether the two participants receive directed Actor and Actee roles or two Mutual roles. The Access data use NULL, N, Y, X, lowercase x, and potentially questionable yes/no forms rather than a single canonical encoding.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_two_sided_flag,&lt;br /&gt;
       quote_nullable(pg_two_sided_flag) AS quoted_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY pg_two_sided_flag&lt;br /&gt;
 ORDER BY row_count DESC,&lt;br /&gt;
          pg_two_sided_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to N in the clean schema. Normalize Y, Y?, X, and lowercase x to Y. During loading, Y creates two Mutual roles; N creates an Actor role for pg_b_actor_id and an Actee role for pg_b_recipient_id. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#126) PANTGRUNT_EVENT multiple-participant flags use noncanonical encodings ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
PANTGRUNTS.MultiActors and PANTGRUNTS.MultiRecipients are required Boolean values, while the corresponding Access fields are nullable text flags. The local conversion data include a lowercase x in pg_multiple_recipient_flag, and future dumps may include questionable yes/no forms.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT flag_name,&lt;br /&gt;
       flag_value,&lt;br /&gt;
       COUNT(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_multiple_actor_flag&amp;#039;, pg_multiple_actor_flag),&lt;br /&gt;
                (&amp;#039;pg_multiple_recipient_flag&amp;#039;, pg_multiple_recipient_flag)&lt;br /&gt;
       ) AS flags(flag_name, flag_value)&lt;br /&gt;
 GROUP BY flag_name,&lt;br /&gt;
          flag_value&lt;br /&gt;
 ORDER BY flag_name,&lt;br /&gt;
          row_count DESC,&lt;br /&gt;
          flag_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize NULL, the empty string, N, N?, and ? to the empty string in the clean schema and load them as false. Normalize Y, Y?, X, and lowercase x to X and load them as true. The sanity check stops conversion if an unrecognized value remains.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#127) PANTGRUNT_EVENT participants are missing from BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Every pantgrunt actor and recipient must be represented by a BIOGRAPHY_DATA row before a corresponding ROLES row can be created. In the local conversion data examined on 2026-08-21, 165 rows had an actor missing from BIOGRAPHY_DATA and 609 rows had a recipient missing from BIOGRAPHY_DATA. A row can be counted in both categories.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(participant.animid))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude any row whose trimmed actor or recipient does not have a matching BIOGRAPHY_DATA animal. Report actor and recipient counts separately in the sanity check. The project investigators must identify the intended individuals and correct the source records in Access.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#128) PANTGRUNT_EVENT focal/date keys have no corresponding follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A pantgrunt event must relate to a WATCHES row. The normal B-record watches are created from clean.follow, but many pantgrunt focal/date keys have no matching follow. The local conversion data examined on 2026-08-21 contained 2,443 normalized focal/date keys without a B-record watch before other pantgrunt exclusions were applied.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg.pg_date,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow AS f&lt;br /&gt;
          WHERE f.fol_b_animid = BTRIM(pg.pg_fol_b_foc_id)&lt;br /&gt;
                AND f.fol_date = pg.pg_date)&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg.pg_date&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Prefer and reuse an existing B-record WATCHES row having the same focal and date. If no such watch exists for an otherwise eligible pantgrunt, reuse an existing Other watch or create a WATCHES row with Type Other. Use the pantgrunt community for a newly created watch. This avoids falsely representing an ad-hoc observation as a follow and avoids the missing-arrival warnings associated with B-record watches. The source association can be corrected later without discarding the converted pantgrunt.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#129) PANTGRUNT_EVENT rows have no focal animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some pantgrunt records are ad-hoc observations and contain a NULL, empty, or whitespace-only pg_fol_b_foc_id. A WATCHES row requires an animal ID, but assigning the unknown individual UNK would incorrectly imply that a focal existed and merely could not be identified. In the local conversion data examined on 2026-08-21, 5,136 rows had no focal value.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
For WATCHES resolution only, normalize an absent pantgrunt focal to the established no-focal AnimID NONE. Create or reuse an Other watch for NONE and the pantgrunt date. Leave the original blank focal value unchanged in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#130) PANTGRUNT_EVENT focal/date keys contain conflicting communities ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Only one Other WATCHES row can exist for a given focal and date, but some pantgrunt keys contain more than one pantgrunt community. After absent focals are normalized to NONE for watch resolution, the local conversion data examined on 2026-08-21 contain 15 conflicting keys: three named-focal keys and twelve additional blank-focal keys that collapse to NONE.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;) AS watch_animid,&lt;br /&gt;
       pg_date,&lt;br /&gt;
       ARRAY_AGG(DISTINCT BTRIM(pg_cl_community_id) ORDER BY BTRIM(pg_cl_community_id)) AS communities,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 GROUP BY COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;),&lt;br /&gt;
          pg_date&lt;br /&gt;
HAVING COUNT(DISTINCT BTRIM(pg_cl_community_id)) &amp;gt; 1&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          watch_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude all pantgrunt rows belonging to a conflicting normalized focal/date key and report the number of affected keys in the sanity check. The project investigators must determine the correct focal, date, or community in Access. The conversion must not choose one community arbitrarily.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#131) PANTGRUNT_EVENT observer values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Status: Unresolved; intentionally retained only in the clean schema.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The pg_observer field is populated only with JF in the local conversion data examined on 2026-08-21, affecting 1,696 rows across 221 normalized watch keys. PANTGRUNTS has no observer column. FOLLOW_OBSERVERS describes observers assigned to follow periods, so creating FOLLOW_OBSERVERS rows from pantgrunt records would require deciding whether pg_observer represents the same concept and period.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_observer,&lt;br /&gt;
       COUNT(*) AS pantgrunt_rows,&lt;br /&gt;
       COUNT(DISTINCT (COALESCE(NULLIF(BTRIM(pg_fol_b_foc_id), &amp;#039;&amp;#039;), &amp;#039;NONE&amp;#039;), pg_date)) AS watch_keys&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_observer IS NOT NULL&lt;br /&gt;
 GROUP BY pg_observer&lt;br /&gt;
 ORDER BY pg_observer;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not convert pg_observer into FOLLOW_OBSERVERS or another production table without an investigator-approved semantic mapping. Retain the values in clean.pantgrunt_event for later review. The pantgrunt loader explicitly documents that this field is omitted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#132) PANTGRUNT_EVENT participants occur outside their biography intervals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
ROLES requires an event participant&amp;#039;s date to fall between BIOGRAPHY_DATA.EntryDate and BIOGRAPHY_DATA.DepartDate, inclusive. Some pantgrunt actors or recipients have a biography row but occur before entry or after departure. In the local conversion data examined on 2026-08-21, 792 rows involved at least one participant outside this interval.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*,&lt;br /&gt;
       participant.source_column,&lt;br /&gt;
       participant.animid,&lt;br /&gt;
       b.b_entrydate,&lt;br /&gt;
       b.b_departdate,&lt;br /&gt;
       CASE&lt;br /&gt;
         WHEN pg.pg_date &amp;lt; b.b_entrydate THEN &amp;#039;BEFORE ENTRY&amp;#039;&lt;br /&gt;
         WHEN pg.pg_date &amp;gt; b.b_departdate THEN &amp;#039;AFTER DEPARTURE&amp;#039;&lt;br /&gt;
       END AS date_problem&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;pg_b_actor_id&amp;#039;, pg.pg_b_actor_id),&lt;br /&gt;
                (&amp;#039;pg_b_recipient_id&amp;#039;, pg.pg_b_recipient_id)&lt;br /&gt;
       ) AS participant(source_column, animid)&lt;br /&gt;
       JOIN clean.biography AS b&lt;br /&gt;
         ON b.b_animid = BTRIM(participant.animid)&lt;br /&gt;
 WHERE pg.pg_date NOT BETWEEN b.b_entrydate AND b.b_departdate&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          participant.source_column,&lt;br /&gt;
          participant.animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude affected rows and report their count in the pantgrunt sanity check. The project investigators must determine whether the pantgrunt date, participant identity, or biography interval should be corrected in Access. Do not weaken the ROLES temporal constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#133) PANTGRUNT_EVENT contains invalid nonempty focal animal IDs ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An Other WATCHES row created for a pantgrunt requires a valid BIOGRAPHY_DATA animal ID. Some pg_fol_b_foc_id values are nonempty but do not match a biography animal and therefore cannot be treated as the no-focal value NONE. In the local conversion data examined on 2026-08-21, 70 rows contained such focal values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg.*&lt;br /&gt;
  FROM clean.pantgrunt_event AS pg&lt;br /&gt;
 WHERE NULLIF(BTRIM(pg.pg_fol_b_foc_id), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
       AND NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.biography AS b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(pg.pg_fol_b_foc_id))&lt;br /&gt;
 ORDER BY pg.pg_date,&lt;br /&gt;
          pg.pg_fol_b_foc_id,&lt;br /&gt;
          pg.pg_time,&lt;br /&gt;
          pg.pg_b_actor_id,&lt;br /&gt;
          pg.pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows and report their count in the sanity check. The project investigators must identify whether each value is a malformed focal ID, descriptive text, or evidence that no focal existed, then correct Access accordingly. Do not automatically convert a nonempty value to NONE or UNK.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#134) PANTGRUNT_EVENT notes may be NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
EVENTS.Notes is required but permits the empty string. The Access pantgrunt table uses NULL for absent notes. In the local conversion data examined on 2026-08-21, 7,339 rows had NULL pg_notes and no rows contained whitespace-only notes.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_notes IS NULL&lt;br /&gt;
       OR (pg_notes &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND BTRIM(pg_notes) = &amp;#039;&amp;#039;)&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
At the production load boundary, convert a NULL or whitespace-only pg_notes value to the empty string and otherwise preserve the source text unchanged. Retain the original value in clean.pantgrunt_event.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#135) PANTGRUNT_EVENT year values have no lossless production destination ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The production pantgrunt model obtains the event date through WATCHES and has no separate year column. Although pg_year appears redundant with pg_date, the local conversion data examined on 2026-08-21 contain 178 rows where pg_year differs from the year extracted from pg_date. Discarding the field as derived data would therefore lose a recorded discrepancy.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;records&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS NULL&lt;br /&gt;
       OR pg_year IS DISTINCT FROM extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_date,&lt;br /&gt;
          pg_fol_b_foc_id,&lt;br /&gt;
          pg_time,&lt;br /&gt;
          pg_b_actor_id,&lt;br /&gt;
          pg_b_recipient_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;summary&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT pg_year,&lt;br /&gt;
       extract(year FROM pg_date)::INTEGER AS date_year,&lt;br /&gt;
       pg_year - extract(year FROM pg_date)::INTEGER AS year_difference,&lt;br /&gt;
       count(*) AS row_count&lt;br /&gt;
  FROM clean.pantgrunt_event&lt;br /&gt;
 WHERE pg_year IS DISTINCT FROM&lt;br /&gt;
         extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 GROUP BY pg_year,&lt;br /&gt;
          extract(year FROM pg_date)::INTEGER&lt;br /&gt;
 ORDER BY pg_year,&lt;br /&gt;
          date_year;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Do not load pg_year into another production column or silently replace pg_date. Retain pg_year in clean.pantgrunt_event for investigator review. The pantgrunt loader explicitly documents that the field is omitted pending a project decision about the 178 discrepancies.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=808</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=808"/>
		<updated>2026-08-21T22:27:29Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark Problem #112 as not resolved.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
ICG or ELO to please confirm.&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=807</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=807"/>
		<updated>2026-08-21T21:33:24Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Problem #112 GROOM_BOUT uses U as an unknown initiator or terminator animal ID&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#112) GROOM_BOUT uses U as an unknown initiator or terminator animal ID ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Some GROOM_BOUT rows contain U in GRM_B_initiator_AnimId or GRM_B_terminator_AnimId. Although U is a valid value for GRM_direction, it is not a valid animal ID.&lt;br /&gt;
&lt;br /&gt;
The production GROOMINGS.Initiator and GROOMINGS.Terminator columns reference participant rows in ROLES. During conversion, the loader creates roles for the focal and partner animals and then looks up the role corresponding to the recorded initiator or terminator. A value of U cannot match either participant, causing the loader to fail with query returned no rows.&lt;br /&gt;
&lt;br /&gt;
This problem became visible after the Problem #99 and Problem #100 exclusions were removed. Those exclusions had prevented many affected grooming bouts from reaching the initiator and terminator lookup.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The following query identifies every occurrence of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; in the initiator and terminator animal-ID fields:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gb.grm_fol_date,&lt;br /&gt;
       gb.grm_fol_b_focal_animid,&lt;br /&gt;
       gb.grm_time_begin,&lt;br /&gt;
       gb.grm_b_partner_animid,&lt;br /&gt;
       gb.grm_direction,&lt;br /&gt;
       invalid_participant.source_column,&lt;br /&gt;
       invalid_participant.animid AS offending_animid,&lt;br /&gt;
       gb.grm_extracted_by,&lt;br /&gt;
       gb.grm_problems,&lt;br /&gt;
       gb.grm_comments&lt;br /&gt;
  FROM clean.groom_bout AS gb&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;GRM_B_initiator_AnimId&amp;#039;, gb.grm_b_initiator_animid),&lt;br /&gt;
                (&amp;#039;GRM_B_terminator_AnimId&amp;#039;, gb.grm_b_terminator_animid)&lt;br /&gt;
       ) AS invalid_participant(source_column, animid)&lt;br /&gt;
 WHERE BTRIM(invalid_participant.animid) = &amp;#039;U&amp;#039;&lt;br /&gt;
 ORDER BY gb.grm_fol_date,&lt;br /&gt;
          gb.grm_fol_b_focal_animid,&lt;br /&gt;
          gb.grm_time_begin,&lt;br /&gt;
          gb.grm_b_partner_animid,&lt;br /&gt;
          invalid_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;proposed&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
During construction of the clean schema, convert U in groom_bout.grm_b_initiator_animid and groom_bout.grm_b_terminator_animid to SQL NULL. Do not change groom_bout.grm_direction when it contains U, because that is a valid direction code representing unknown grooming direction.&lt;br /&gt;
&lt;br /&gt;
The grooming loader already treats a NULL initiator or terminator as unknown and stores NULL in the corresponding production column without attempting to find a participant role.&lt;br /&gt;
&lt;br /&gt;
Other nonempty initiator or terminator animal IDs that do not identify either the focal or partner animal should remain conversion errors and be reviewed separately rather than being converted automatically to NULL.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=806</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=806"/>
		<updated>2026-08-21T12:07:19Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: fix non-MW formatting in 114-117&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access BRECORD_NOTES.BREC_time column is stored as a timestamp and is converted to a PostgreSQL TIME value in the tidy schema. Problem #10 corrected records whose timestamp had the wrong date component, but converting the datatype preserves seconds. The production EVENTS.Start and EVENTS.Stop columns record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded. The production schema previously reserved the sdb_no_time midnight sentinel for aggression events, so it rejected otherwise-loadable B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time. Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from BREC_time in tidy_cleanups.sql before the column is converted to TIME. Do not round the values. Convert a SQL NULL time to sdb_no_time at the production load boundary and preserve an existing midnight value as that sentinel. Permit both aggression and B-record-note EVENTS rows to use sdb_no_time, with both Start and Stop set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production BRECORD_NOTES row belongs to an EVENTS row, which in turn must belong to a WATCHES row. The ordinary B-record watches are created from clean.follow. Some B-record notes have no follow with the same focal individual and date. It is not yet known which records truly lack a follow and which contain an incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or date and would make the affected records difficult to identify for later review. In the local conversion data examined on 2026-08-20, 35,517 of 603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion. Retain them unchanged in clean.brecord_notes so that the focal and date can be reviewed and corrected in Access.  The loader must reuse only an existing type-B WATCHES row having the same focal and date; it must not create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the sdb_no_time sentinel, production event times must be between sdb_min_event_start (04:00) and sdb_max_event_stop (20:00), inclusive. Some B-record-note times are outside those limits and their correct values must be determined from the source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an out-of-range time other than midnight. Of those, 711 had a matching follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while retaining them in clean.brecord_notes. Correct the source times in Access after reviewing the original records. Do not clamp the values or weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production BRECORD_NOTES text columns are NOT NULL. Observation must also be trimmed of leading and trailing spaces. The remaining text columns permit the empty string to represent absent text but reject values consisting only of whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some whitespace-only values. In the local conversion data examined on 2026-08-20, the rows otherwise eligible for conversion included 41 NULL observations and 14,349 untrimmed observations. The optional columns contained between 15,080 and 566,253 NULL values. Four comments, three Voc values, and three VocID values consisted only of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load boundary. Convert a NULL or whitespace-only value to the empty string. Trim Observation because its production column requires it, but preserve leading and trailing spaces in other nonempty source text. Keep the unchanged source-like values available in clean.brecord_notes.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=805</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=805"/>
		<updated>2026-08-21T00:43:37Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Add problems #114 - #117 regarding brecord notes&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#114) BRECORD_NOTES times are not uniformly compatible with EVENTS ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The Access &amp;lt;pre&amp;gt;BRECORD_NOTES.BREC_time&amp;lt;/pre&amp;gt; column is stored as a&lt;br /&gt;
timestamp and is converted to a PostgreSQL &amp;lt;pre&amp;gt;TIME&amp;lt;/pre&amp;gt; value in the&lt;br /&gt;
&amp;lt;pre&amp;gt;tidy&amp;lt;/pre&amp;gt; schema.  Problem #10 corrected records whose timestamp had&lt;br /&gt;
the wrong date component, but converting the datatype preserves seconds.&lt;br /&gt;
The production &amp;lt;pre&amp;gt;EVENTS.Start&amp;lt;/pre&amp;gt; and &amp;lt;pre&amp;gt;EVENTS.Stop&amp;lt;/pre&amp;gt; columns&lt;br /&gt;
record times to the minute and reject nonzero seconds.&lt;br /&gt;
&lt;br /&gt;
Some B-record notes also use midnight to indicate that no time was recorded.&lt;br /&gt;
The production schema previously reserved the &amp;lt;pre&amp;gt;sdb_no_time&amp;lt;/pre&amp;gt;&lt;br /&gt;
midnight sentinel for aggression events, so it rejected otherwise-loadable&lt;br /&gt;
B-record notes with no recorded time.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 22 rows contained&lt;br /&gt;
nonzero seconds, 46 rows contained midnight, and no rows had a SQL NULL time.&lt;br /&gt;
Forty-five of the midnight rows had a matching follow-derived B-record watch.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       extract(second FROM brec_time) AS seconds&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NULL&lt;br /&gt;
       OR brec_time = &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       OR extract(second FROM brec_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Truncate nonzero seconds from &amp;lt;pre&amp;gt;BREC_time&amp;lt;/pre&amp;gt; in&lt;br /&gt;
&amp;lt;pre&amp;gt;tidy_cleanups.sql&amp;lt;/pre&amp;gt; before the column is converted to&lt;br /&gt;
&amp;lt;pre&amp;gt;TIME&amp;lt;/pre&amp;gt;.  Do not round the values.  Convert a SQL NULL time to&lt;br /&gt;
&amp;lt;pre&amp;gt;sdb_no_time&amp;lt;/pre&amp;gt; at the production load boundary and preserve an&lt;br /&gt;
existing midnight value as that sentinel.  Permit both aggression and&lt;br /&gt;
B-record-note &amp;lt;pre&amp;gt;EVENTS&amp;lt;/pre&amp;gt; rows to use &amp;lt;pre&amp;gt;sdb_no_time&amp;lt;/pre&amp;gt;, with&lt;br /&gt;
both &amp;lt;pre&amp;gt;Start&amp;lt;/pre&amp;gt; and &amp;lt;pre&amp;gt;Stop&amp;lt;/pre&amp;gt; set to the sentinel.&lt;br /&gt;
&lt;br /&gt;
== * (#115) BRECORD_NOTES rows lack a matching follow-derived WATCHES row ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
A production &amp;lt;pre&amp;gt;BRECORD_NOTES&amp;lt;/pre&amp;gt; row belongs to an&lt;br /&gt;
&amp;lt;pre&amp;gt;EVENTS&amp;lt;/pre&amp;gt; row, which in turn must belong to a &amp;lt;pre&amp;gt;WATCHES&amp;lt;/pre&amp;gt;&lt;br /&gt;
row.  The ordinary B-record watches are created from &amp;lt;pre&amp;gt;clean.follow&amp;lt;/pre&amp;gt;.&lt;br /&gt;
Some B-record notes have no follow with the same focal individual and date.&lt;br /&gt;
It is not yet known which records truly lack a follow and which contain an&lt;br /&gt;
incorrect focal or date.&lt;br /&gt;
&lt;br /&gt;
Creating provisional watches would require guessing the focal, community, or&lt;br /&gt;
date and would make the affected records difficult to identify for later&lt;br /&gt;
review.  In the local conversion data examined on 2026-08-20, 35,517 of&lt;br /&gt;
603,977 B-record-note rows lacked a matching follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
          WHERE follow.fol_b_animid =&lt;br /&gt;
                  brecord_notes.brec_fol_b_focal_animid&lt;br /&gt;
                AND follow.fol_date = brecord_notes.brec_fol_date)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude these rows from the production conversion.  Retain them&lt;br /&gt;
unchanged in &amp;lt;pre&amp;gt;clean.brecord_notes&amp;lt;/pre&amp;gt; so that the focal and date can be&lt;br /&gt;
reviewed and corrected in Access.  The loader must reuse only an existing&lt;br /&gt;
type-B &amp;lt;pre&amp;gt;WATCHES&amp;lt;/pre&amp;gt; row having the same focal and date; it must not&lt;br /&gt;
create a watch for an unmatched B-record note.&lt;br /&gt;
&lt;br /&gt;
== * (#116) BRECORD_NOTES contains times outside the EVENTS limits ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Except for the &amp;lt;pre&amp;gt;sdb_no_time&amp;lt;/pre&amp;gt; sentinel, production event times must&lt;br /&gt;
be between &amp;lt;pre&amp;gt;sdb_min_event_start&amp;lt;/pre&amp;gt; (04:00) and&lt;br /&gt;
&amp;lt;pre&amp;gt;sdb_max_event_stop&amp;lt;/pre&amp;gt; (20:00), inclusive.  Some B-record-note times&lt;br /&gt;
are outside those limits and their correct values must be determined from the&lt;br /&gt;
source records.&lt;br /&gt;
&lt;br /&gt;
In the local conversion data examined on 2026-08-20, 751 rows had an&lt;br /&gt;
out-of-range time other than midnight.  Of those, 711 had a matching&lt;br /&gt;
follow-derived watch and would otherwise be loadable.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brecord_notes.*&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
 WHERE brec_time IS NOT NULL&lt;br /&gt;
       AND brec_time &amp;lt;&amp;gt; &amp;#039;00:00&amp;#039;::TIME&lt;br /&gt;
       AND (brec_time &amp;lt; &amp;#039;04:00&amp;#039;::TIME&lt;br /&gt;
            OR brec_time &amp;gt; &amp;#039;20:00&amp;#039;::TIME)&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Temporarily exclude the offending rows from the production conversion while&lt;br /&gt;
retaining them in &amp;lt;pre&amp;gt;clean.brecord_notes&amp;lt;/pre&amp;gt;.  Correct the source times&lt;br /&gt;
in Access after reviewing the original records.  Do not clamp the values or&lt;br /&gt;
weaken the production time constraints.&lt;br /&gt;
&lt;br /&gt;
== (#117) BRECORD_NOTES text values do not satisfy production constraints ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
All production &amp;lt;pre&amp;gt;BRECORD_NOTES&amp;lt;/pre&amp;gt; text columns are&lt;br /&gt;
&amp;lt;pre&amp;gt;NOT NULL&amp;lt;/pre&amp;gt;.  &amp;lt;pre&amp;gt;Observation&amp;lt;/pre&amp;gt; must also be trimmed of&lt;br /&gt;
leading and trailing spaces.  The remaining text columns permit the empty&lt;br /&gt;
string to represent absent text but reject values consisting only of&lt;br /&gt;
whitespace.&lt;br /&gt;
&lt;br /&gt;
The Access source uses SQL NULL for absent text and contains some&lt;br /&gt;
whitespace-only values.  In the local conversion data examined on&lt;br /&gt;
2026-08-20, the rows otherwise eligible for conversion included 41 NULL&lt;br /&gt;
observations and 14,349 untrimmed observations.  The optional columns&lt;br /&gt;
contained between 15,080 and 566,253 NULL values.  Four comments, three&lt;br /&gt;
&amp;lt;pre&amp;gt;Voc&amp;lt;/pre&amp;gt; values, and three &amp;lt;pre&amp;gt;VocID&amp;lt;/pre&amp;gt; values consisted only&lt;br /&gt;
of whitespace.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT brec_fol_date,&lt;br /&gt;
       brec_fol_b_focal_animid,&lt;br /&gt;
       brec_time,&lt;br /&gt;
       field_name,&lt;br /&gt;
       quote_nullable(field_value) AS field_value&lt;br /&gt;
  FROM clean.brecord_notes&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Observation&amp;#039;, brec_observation),&lt;br /&gt;
                (&amp;#039;Comments&amp;#039;, brec_comments),&lt;br /&gt;
                (&amp;#039;Observer&amp;#039;, brec_observer),&lt;br /&gt;
                (&amp;#039;Translator&amp;#039;, brec_translator),&lt;br /&gt;
                (&amp;#039;TranscribedBy&amp;#039;, brec_transcribed_by),&lt;br /&gt;
                (&amp;#039;Voc&amp;#039;, brec_voc),&lt;br /&gt;
                (&amp;#039;VocID&amp;#039;, brec_vocid),&lt;br /&gt;
                (&amp;#039;GroomingAggression&amp;#039;, brec_grooming_aggressionflag),&lt;br /&gt;
                (&amp;#039;Duplicate&amp;#039;, brec_duplicate_flag)&lt;br /&gt;
       ) AS source_text(field_name, field_value)&lt;br /&gt;
 WHERE field_value IS NULL&lt;br /&gt;
       OR (field_value &amp;lt;&amp;gt; &amp;#039;&amp;#039; AND btrim(field_value) = &amp;#039;&amp;#039;)&lt;br /&gt;
       OR (field_name = &amp;#039;Observation&amp;#039;&lt;br /&gt;
           AND field_value &amp;lt;&amp;gt; btrim(field_value))&lt;br /&gt;
 ORDER BY brec_fol_date,&lt;br /&gt;
          brec_fol_b_focal_animid,&lt;br /&gt;
          brec_time,&lt;br /&gt;
          field_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Normalize only what the production constraints require at the final load&lt;br /&gt;
boundary.  Convert a NULL or whitespace-only value to the empty string.&lt;br /&gt;
Trim &amp;lt;pre&amp;gt;Observation&amp;lt;/pre&amp;gt; because its production column requires it, but&lt;br /&gt;
preserve leading and trailing spaces in other nonempty source text.  Keep the&lt;br /&gt;
unchanged source-like values available in &amp;lt;pre&amp;gt;clean.brecord_notes&amp;lt;/pre&amp;gt;.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=804</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=804"/>
		<updated>2026-08-20T18:36:56Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: fix #71 processing note list formatting&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=803</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=803"/>
		<updated>2026-08-20T18:36:09Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: expand #71 processing note&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of&lt;br /&gt;
Problem #70. As such, there is not a separate exclusion clause in the conversion&lt;br /&gt;
process for these data. As a corollary, addressing all issues in Problem #70&lt;br /&gt;
would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
Problem #70 concerns aggression rows without a matching follow for the same&lt;br /&gt;
focal ID and date. Canonical master solves that relationship problem by creating&lt;br /&gt;
a B-record WATCHES row when no suitable watch exists. An aggression therefore&lt;br /&gt;
does not inherently require a FOLLOW.&lt;br /&gt;
&lt;br /&gt;
Problem #71 is a stricter subset: its focal IDs never occur among follow focal&lt;br /&gt;
IDs. More importantly, a read-only check of the local data found that the&lt;br /&gt;
documented 16 IDs are also absent from clean.biography, including values such as&lt;br /&gt;
DL DUK, GA GGL, GROUP, Males, and stranger.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;WATCHES.AnimID → BIOGRAPHY_DATA.AnimID&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Consequently:&lt;br /&gt;
&lt;br /&gt;
- If a #71 focal ID is a valid biography ID, the #70 WATCHES solution handles&lt;br /&gt;
  it.&lt;br /&gt;
- If it is not a valid biography ID, creating the WATCHES row fails its foreign&lt;br /&gt;
  key.&lt;br /&gt;
- The documented 16 IDs appear to be in the latter category and therefore&lt;br /&gt;
  require a separate source-data decision or clean-stage mapping.&lt;br /&gt;
&lt;br /&gt;
There is also a branch-specific complication: the temporary branch loader still&lt;br /&gt;
has the old #70 WHERE EXISTS filter. It excludes all rows lacking a matching&lt;br /&gt;
follow, including every #71 row. Thus the successful mung conversion does not&lt;br /&gt;
demonstrate that these rows were converted. Canonical master does not have that&lt;br /&gt;
filter and would expose invalid #71 focal IDs when attempting to create their&lt;br /&gt;
WATCHES rows.&lt;br /&gt;
&lt;br /&gt;
Therefore, the Problem #71 processing note is too broad. The #70 solution&lt;br /&gt;
removes the requirement for a corresponding follow, but it does not resolve&lt;br /&gt;
invalid or composite focal animal IDs. Problem #71 should remain a separate&lt;br /&gt;
unresolved data-quality problem, and mung’s stale #70 exclusion should&lt;br /&gt;
eventually be removed when testing the full solution.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=802</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=802"/>
		<updated>2026-08-20T18:27:27Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark #77 as resolved.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=801</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=801"/>
		<updated>2026-08-20T16:58:35Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark #82 as resolved and add context regarding the resolution&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
Added `NODATA` to `AGG_SEVERITIES` for aggression records where no&lt;br /&gt;
fight-category data were recorded. This is distinct from `unrated`, which means&lt;br /&gt;
the aggression was not evaluated for severity.&lt;br /&gt;
&lt;br /&gt;
During construction of the `clean` schema, `ae_fight_category` values are&lt;br /&gt;
trimmed. `NULL`, empty-string, and whitespace-only values are converted to&lt;br /&gt;
`NODATA`. The aggression loader then copies the cleaned value directly into&lt;br /&gt;
`AGGRESSIONS.Severity`, allowing unexpected nonblank codes to remain visible as&lt;br /&gt;
conversion errors.&lt;br /&gt;
&lt;br /&gt;
Implemented in commit `d911ef2` (`Resolve Problem #82 with NODATA severity`). A&lt;br /&gt;
complete conversion using the updated code succeeded.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=800</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=800"/>
		<updated>2026-08-20T13:45:18Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark #81 as resolved and add resolution context&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
=== resolution ===&lt;br /&gt;
&lt;br /&gt;
The schema already has a deliberate non-NULL representation for an unknown&lt;br /&gt;
aggression time: 00:00, defined as sdb_no_time.&lt;br /&gt;
&lt;br /&gt;
The design is explicit:&lt;br /&gt;
&lt;br /&gt;
- EVENTS.Start and EVENTS.Stop remain NOT NULL.&lt;br /&gt;
- Normal event times must be between 04:00 and 20:00.&lt;br /&gt;
- Aggressions alone may use 00:00 to mean “no time was recorded.”&lt;br /&gt;
- Both start and stop must be 00:00 together.&lt;br /&gt;
&lt;br /&gt;
Existing implementation Commit 37bb5493b7571084e6d57aa45c5440dcb0db357e Karl O.&lt;br /&gt;
Pinc 2026-07-15&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=799</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=799"/>
		<updated>2026-08-20T01:49:31Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark #80 as resolved.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=798</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=798"/>
		<updated>2026-08-20T01:47:24Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark #78 as resolved.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=797</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=797"/>
		<updated>2026-08-20T01:39:51Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark #73 as resolved.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=796</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=796"/>
		<updated>2026-08-20T01:31:23Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark #70 as resolved&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=795</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=795"/>
		<updated>2026-08-20T01:30:56Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Update solution (--&amp;gt; resolution) to #70&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Resolution ===&lt;br /&gt;
&lt;br /&gt;
Resolved by commit `1cc9e694a1f6a53c396d9074f77d5d279018721b`.&lt;br /&gt;
&lt;br /&gt;
A matching FOLLOW row is not required for an aggression event. WATCHES&lt;br /&gt;
represents the observation context independently of FOLLOW. A B-record&lt;br /&gt;
WATCHES row may represent either a follow or an ad-hoc observation such as&lt;br /&gt;
an aggression.&lt;br /&gt;
&lt;br /&gt;
During conversion, `load_aggressions.m4` looks for an existing B-record&lt;br /&gt;
WATCHES row having the aggression event&amp;#039;s focal individual and date. If no&lt;br /&gt;
such row exists, the loader creates one using the aggression event&amp;#039;s focal,&lt;br /&gt;
community, and date. The aggression&amp;#039;s EVENTS row is then related to that&lt;br /&gt;
WATCHES row.&lt;br /&gt;
&lt;br /&gt;
The original diagnostic query will continue to identify aggression records&lt;br /&gt;
without matching FOLLOW rows. These rows are not necessarily bad data and&lt;br /&gt;
are not expected to disappear.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=794</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=794"/>
		<updated>2026-08-19T20:30:41Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark #84 as resolved.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=793</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=793"/>
		<updated>2026-08-19T20:29:20Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark #63 as resolved&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=792</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=792"/>
		<updated>2026-08-19T20:23:18Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark #13 and #14 as resolved&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=791</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=791"/>
		<updated>2026-08-19T20:05:09Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Fix heading level on #97 solution&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=790</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=790"/>
		<updated>2026-08-19T19:39:08Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark #49 as not resolved and update query to show age at time of follow&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate)&lt;br /&gt;
         - INTERVAL &amp;#039;15 years&amp;#039; AS older_than_limit_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate&lt;br /&gt;
             &amp;lt;= fa.fa_fol_date&lt;br /&gt;
                - INTERVAL &amp;#039;1 year&amp;#039;&lt;br /&gt;
                - INTERVAL &amp;#039;14 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: Improved but there are still 48 offending records&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==SOLUTION==&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=789</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=789"/>
		<updated>2026-08-19T19:36:08Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: update #48 row count&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;86&amp;lt;/s&amp;gt; 48 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==SOLUTION==&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=788</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=788"/>
		<updated>2026-08-19T19:34:42Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark #47 as not resolved and modify query to calculate age at time of follow.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT fa.*&lt;br /&gt;
     , b.b_sex&lt;br /&gt;
     , b.b_birthdate&lt;br /&gt;
     , AGE(fa.fa_fol_date, b.b_birthdate) AS age_at_follow&lt;br /&gt;
     , INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
         - AGE(fa.fa_fol_date, b.b_birthdate) AS younger_than_five_by&lt;br /&gt;
  FROM clean.follow_arrival AS fa&lt;br /&gt;
  JOIN clean.biography AS b&lt;br /&gt;
    ON b.b_animid = fa.fa_b_arr_animid&lt;br /&gt;
 WHERE b.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
       AND fa.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
       AND b.b_birthdate &amp;gt; fa.fa_fol_date - INTERVAL &amp;#039;5 years&amp;#039;&lt;br /&gt;
 ORDER BY fa.fa_fol_date&lt;br /&gt;
        , fa.fa_fol_b_focal_animid&lt;br /&gt;
        , fa.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-19: improved but there are still 12 offending records&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==SOLUTION==&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=787</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=787"/>
		<updated>2026-08-19T02:02:29Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Problem #113 The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
ICG fixed in Access 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN fixed 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====SOLUTION====&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==SOLUTION==&lt;br /&gt;
ICG FIXED IN ACCESS 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
ICG CHANGED ALL NULLS TO &amp;#039;UNK&amp;#039; IN ACCESS. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
&lt;br /&gt;
HAVE A STUDENT CHECK AGAINST PAPER&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#113) The Access BIOGRAPHY.NONE row lacks values required by BIOGRAPHY_DATA ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The A-record groom-scan conversion represents feeding-station scans as A-record WATCHES. These scans have no focal individual, so the conversion uses the special AnimID &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; for the related WATCHES row rather than incorrectly identifying the focal individual as unknown with &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;. The production BIOGRAPHY_DATA table must therefore contain a &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; row.&lt;br /&gt;
&lt;br /&gt;
The conversion originally created this row in &amp;lt;code&amp;gt;clean.BIOGRAPHY&amp;lt;/code&amp;gt;. After the row was added to the Access BIOGRAPHY table, retaining that INSERT caused the clean-schema build to fail with a duplicate &amp;lt;code&amp;gt;BIOGRAPHY&amp;lt;/code&amp;gt; primary key for &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. Removing the INSERT exposed a second problem: the Access row contains NULL in fields that are required by BIOGRAPHY_DATA, beginning with &amp;lt;code&amp;gt;BCCertainty&amp;lt;/code&amp;gt;, so it cannot be loaded directly into the production table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
The affected row is the single Access BIOGRAPHY record whose &amp;lt;code&amp;gt;B_AnimID&amp;lt;/code&amp;gt; is &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt;. A diagnostic query is unnecessary because this is an intentionally defined sentinel record rather than an unidentified set of source-data rows.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Keep Access authoritative for the existence of the &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; BIOGRAPHY row. During construction of the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, update only that row with the previously established no-focal placeholder values required by BIOGRAPHY_DATA. Do not insert a second row, modify the restored Access data in &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt;, or apply the correction directly to &amp;lt;code&amp;gt;sokwedb&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This preserves the decision that A-record groom scans use &amp;lt;code&amp;gt;NONE&amp;lt;/code&amp;gt; when no focal individual exists while allowing the Access-supplied sentinel record to satisfy the production biography constraints.&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=779</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=779"/>
		<updated>2026-08-18T19:14:08Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Problem #58 update query (OS_FOL_B_focal_AnimId -&amp;gt; OS_FOL_B_focal_AnimID)&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-14: fixed in Access by ICG&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MGF (how to we treat MGF?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
8/14/2026 - ICG fixed in Access - all non-MGF entries changed to &amp;quot;0&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-08-18: ICG fixed in Access. FAR--&amp;gt;FAY, FAC--&amp;gt;FIC&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG FIXED IN ACCESS - ALL WERE LEGIT DUPLICATES. REMOVED. 8/18/2026&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
  GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_FOL_B_focal_AnimID&amp;quot;&lt;br /&gt;
  , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
) AS os&lt;br /&gt;
ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimID&amp;quot; = os.the_animid&lt;br /&gt;
  AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
These are all cases when a juvenile arrival has been extracted from b record notes, but there seems to be no tiki. Do we add WATCHES for these?&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===solution===&lt;br /&gt;
&lt;br /&gt;
Fixed in Access by ICG - duplicated update_rationale&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=770</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=770"/>
		<updated>2026-08-05T23:39:01Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Problem #111 There are GROOM_SCAN_AREC records involving individuals before their entry into the study&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#111) There are GROOM_SCAN_AREC records involving individuals before their entry into the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 130 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is before the individual&amp;#039;s BIOGRAPHY_DATA.EntryDate. The records involve 10 distinct individuals on 21 dates. The offending individual occurs in Chimp_1 in 37 records and in Chimp_2 in 93 records; no record has both participants before their entry dates. After the temporary exclusions for Problems #107 through #110, 30 otherwise-loadable records on 11 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography entry dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.entrydate,&lt;br /&gt;
       bio.entrydate - gs.date AS days_before_entry&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;lt; bio.entrydate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=769</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=769"/>
		<updated>2026-08-05T23:24:48Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Problem #110 There are GROOM_SCAN_AREC records involving individuals after their departure from the study&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#110) There are GROOM_SCAN_AREC records involving individuals after their departure from the study. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,090 GROOM_SCAN_AREC records where at least one participant&amp;#039;s scan date is after the individual&amp;#039;s BIOGRAPHY_DATA.DepartDate. The records involve 18 distinct individuals on 177 dates. The offending individual occurs in Chimp_1 in 1,928 records and in Chimp_2 in 162 records; no record has both participants after their departure dates. After the temporary exclusions for Problems #107, #108, and #109, 1,725 otherwise-loadable records on 147 dates are affected. The production schema requires event participants to be under study on the WATCHES.Date, so these records cannot be loaded until the project investigators determine whether the scan participants or biography departure dates should be corrected.&lt;br /&gt;
&lt;br /&gt;
The trigger error originally reported PT&amp;#039;s EntryDate, 1970-09-11, as its DepartDate due to an error in the trigger&amp;#039;s diagnostic query. PT&amp;#039;s actual BIOGRAPHY_DATA.DepartDate is 1973-04-17, which is before the first affected PT scan on 1973-05-25. The constraint comparison itself used the correct DepartDate and therefore rejected the row correctly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       bad_participant.animid,&lt;br /&gt;
       bio.departdate,&lt;br /&gt;
       gs.date - bio.departdate AS days_after_departure&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
       JOIN sokwedb.biography_data AS bio&lt;br /&gt;
         ON bio.animid = bad_participant.animid&lt;br /&gt;
 WHERE gs.date &amp;gt; bio.departdate&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=768</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=768"/>
		<updated>2026-08-05T23:05:31Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Problem #109 There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#109) There are GROOM_SCAN_AREC records whose dates have no ATTENDANCE records from which to obtain a community. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,111 GROOM_SCAN_AREC records on 134 dates for which there is no ATTENDANCE record on the same date. GROOM_SCAN_AREC does not contain a community value, and the production loader obtains the WATCHES.CommID value from ATTENDANCE.A_CL_Community_ID by date. Consequently, the loader cannot create the required WATCHES row for these scans without assuming a community. After the temporary exclusions for Problems #107 and #108, 4,786 otherwise-loadable records on 133 dates are affected. The project investigators must determine the appropriate community for these scan dates.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM clean.attendance AS attendance&lt;br /&gt;
          WHERE attendance.a_date = gs.date)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=767</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=767"/>
		<updated>2026-08-05T20:35:25Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Problem #108 There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#108) There are GROOM_SCAN_AREC participants that are not in BIOGRAPHY_DATA. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,124 GROOM_SCAN_AREC records having at least one participant that is not a BIOGRAPHY_DATA.AnimID value. There are 81 distinct missing participant values. The missing value occurs in Chimp_1 in 112 records and in Chimp_2 in 1,016 records; four records have missing values in both columns. None of the 81 values are present in clean.BIOGRAPHY, from which BIOGRAPHY_DATA is loaded. The production loader cannot create the related ROLES rows until the project investigators determine whether these values identify individuals that should be added to BIOGRAPHY, corrected to existing AnimID values, or otherwise represented.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader, and sokwedb.BIOGRAPHY_DATA, the referenced production table&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction,&lt;br /&gt;
       bad_participant.source_column,&lt;br /&gt;
       quote_nullable(bad_participant.animid) AS missing_animid,&lt;br /&gt;
       length(bad_participant.animid) AS animid_length&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
       CROSS JOIN LATERAL (&lt;br /&gt;
         VALUES (&amp;#039;Chimp_1&amp;#039;, gs.chimp_1),&lt;br /&gt;
                (&amp;#039;Chimp_2&amp;#039;, gs.chimp_2)&lt;br /&gt;
       ) AS bad_participant(source_column, animid)&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM sokwedb.biography_data AS bio&lt;br /&gt;
          WHERE bio.animid = bad_participant.animid)&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2,&lt;br /&gt;
          bad_participant.source_column;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=766</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=766"/>
		<updated>2026-08-05T20:29:40Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Problem #107 There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#107) There are GROOM_SCAN_AREC records having the same chimpanzee as both grooming participants. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,784 GROOM_SCAN_AREC records where Chimp_1 and Chimp_2 are the same individual. The records span 210 dates and 39 distinct chimpanzees, and all have a Direction value of M. The production loader converts Direction M into two Mutual ROLES rows. These records cannot be loaded because a participant may occur only once in an event and the attendance groom-scan rules require two different participants. The project investigators must determine whether these records represent self-grooming or erroneous participant data and how they should be represented in the production database.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.date,&lt;br /&gt;
       gs.time,&lt;br /&gt;
       gs.chimp_1,&lt;br /&gt;
       gs.chimp_2,&lt;br /&gt;
       gs.direction&lt;br /&gt;
  FROM clean.groom_scan_arec AS gs&lt;br /&gt;
 WHERE gs.chimp_1 = gs.chimp_2&lt;br /&gt;
 ORDER BY gs.date,&lt;br /&gt;
          gs.time,&lt;br /&gt;
          gs.chimp_1,&lt;br /&gt;
          gs.chimp_2;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=765</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=765"/>
		<updated>2026-08-05T02:57:30Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Problem #106 There are GROOM_SCANS records whose extracted-by person is not in PEOPLE.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#106) There are GROOM_SCANS records whose extracted-by person is not in PEOPLE. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 44,679 GROOM_SCANS records where the GS_extracted_by value is not a PEOPLE.Person value. All GROOM_SCANS records are affected. There are four distinct offending values: KASEN, KAREN MCLELLAN, Anika Richter, and Janelle Carmichael.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying clean, the schema read by the production loader&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT gs.gs_date,&lt;br /&gt;
       gs.gs_fol_b_focal_animid,&lt;br /&gt;
       gs.gs_time,&lt;br /&gt;
       gs.gs_b_chimp1_animid,&lt;br /&gt;
       gs.gs_b_chimp2_animid,&lt;br /&gt;
       quote_nullable(gs.gs_extracted_by) AS gs_extracted_by,&lt;br /&gt;
       length(gs.gs_extracted_by) AS extracted_by_length&lt;br /&gt;
  FROM clean.groom_scans AS gs&lt;br /&gt;
 WHERE NOT EXISTS (&lt;br /&gt;
         SELECT 1&lt;br /&gt;
           FROM codes.people AS people&lt;br /&gt;
          WHERE people.person = gs.gs_extracted_by)&lt;br /&gt;
 ORDER BY gs.gs_date,&lt;br /&gt;
          gs.gs_fol_b_focal_animid,&lt;br /&gt;
          gs.gs_time,&lt;br /&gt;
          gs.gs_b_chimp1_animid,&lt;br /&gt;
          gs.gs_b_chimp2_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=764</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=764"/>
		<updated>2026-08-04T20:15:53Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark Problems #103 and #104 as resolved&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit f2ddb85cd22d3e5233d049d92ae8cb361a538114&lt;br /&gt;
&lt;br /&gt;
== (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Addressed with commit 5e0e787825e847eac1cc93375e26ea1d3f715d3e&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=763</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=763"/>
		<updated>2026-07-26T21:32:42Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark Problem #60 as resolved.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Codified in clean per commit b793345f4098fa012d559b8f2d01f06d5418126c.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=762</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=762"/>
		<updated>2026-07-26T20:41:31Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Re-open Problem #79 that has fewer conflicts but is not fully resolved.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== * (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Suggest instead addressing this in clean:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;46&amp;lt;/s&amp;gt; 6 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-26: row count is much improved but there are six records that violate this constraint.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=761</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=761"/>
		<updated>2026-07-26T17:27:19Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Resolve Problems #77, #78, and #80 - changes NULL values to empty strings in clean.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== * (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Suggest instead addressing this in clean:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
Addressed with commit a84462b47c8b2100caa02958623817981179d068.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=760</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=760"/>
		<updated>2026-07-23T22:20:36Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Update Problem #43 row count&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;276&amp;lt;/s&amp;gt; 75 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== * (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Suggest instead addressing this in clean:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=759</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=759"/>
		<updated>2026-07-23T21:16:02Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Mark Problem #54 as solved with caveat of complications with Problem #47&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: The proposed solution has been applied to the clean schema (see commit 06c4178). However, please note that this does not solve Problem #47 and, in fact, forces the fa_type_of_cycle type of records associated with Problem #47 to &amp;#039;U&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== * (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Suggest instead addressing this in clean:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=758</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=758"/>
		<updated>2026-07-23T18:23:33Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: Problem #47 provide summary of possible values&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
option matrix:&amp;lt;br&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! Value !! Meaning !! Accepted under age 5?&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; || Not swollen || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; || Missing data || Yes&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; || Adolescent swelling || No&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;0.25&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;0.75&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;1&amp;lt;/code&amp;gt; || Increasing observed swelling || No—female must be at least 8&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt; || Male/not applicable || No—prohibited for females&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== * (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Suggest instead addressing this in clean:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=757</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=757"/>
		<updated>2026-07-23T18:07:18Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: minor format change to Problem #49&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== * (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Suggest instead addressing this in clean:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
	<entry>
		<id>https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=756</id>
		<title>Conversion Data Issues</title>
		<link rel="alternate" type="text/html" href="https://sokwe.janegoodall.org/w/index.php?title=Conversion_Data_Issues&amp;diff=756"/>
		<updated>2026-07-23T18:04:50Z</updated>

		<summary type="html">&lt;p&gt;StevanEarl: add context to Problem #49&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!-- Using sections, because avoiding any empty lines while using numbered lists is painful.  --&amp;gt;&lt;br /&gt;
&amp;lt;!-- Numbering, so that we can refer to problems by number.&lt;br /&gt;
     Please don&amp;#039;t change the numbering. --&amp;gt;&lt;br /&gt;
This page lists all the problems with the data that were encountered during the data conversion process, and how the issue was resolved.&lt;br /&gt;
&lt;br /&gt;
The problems are numbered, in the order in which they were encountered during the conversion.&lt;br /&gt;
&lt;br /&gt;
Given a choice, earlier problems should be solved before later problems.&lt;br /&gt;
This allows the later steps in the conversion to receive &amp;quot;correct&amp;quot; data, which help eliminate spurious problems, and aids the discovery of problems hidden by bad data.&lt;br /&gt;
&lt;br /&gt;
Unsolved problems are marked with an &amp;lt;code&amp;gt;*&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== (#1) FOLLOW_MAP_TIME duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;FOLLOW_MAP_TIME&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;FMT_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;FMT_time&amp;lt;/code&amp;gt;.&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;FMT_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;FMT_time&amp;quot; AS the_time&lt;br /&gt;
            FROM &amp;quot;FOLLOW_MAP_TIME&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;FMT_FOL_date&amp;quot;, &amp;quot;FMT_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FMT_time&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS fmt&lt;br /&gt;
      ON (&amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_date&amp;quot; = fmt.the_date&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_FOL_B_focal_AnimID&amp;quot; = fmt.the_animid&lt;br /&gt;
          AND &amp;quot;FOLLOW_MAP_TIME&amp;quot;.&amp;quot;FMT_time&amp;quot; = fmt.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access 11/17/23 by ICG. Checked all against Tikis.&lt;br /&gt;
When run in SokweDB, table does not exist&lt;br /&gt;
&lt;br /&gt;
== (#2) SUBADULT_ARRIVALS_LOG has a textual SA_first_tiki_date column ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The SUBADULT_ARRIVALS_LOG table has a column that is supposed to contain a date, but instead contains the string &amp;quot;TEXTY&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
Likely, the entire row is bad.  The row contains:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
       SA_B_AnimID | SA_first_tiki_date | SA_notes&lt;br /&gt;
       -------------+--------------------+----------&lt;br /&gt;
	TEXTY       | TEXTY              | TEXTY&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from  &amp;quot;SUBADULT_ARRIVALS_LOG&amp;quot; where &amp;quot;SA_first_tiki_date&amp;quot; = &amp;#039;TEXTY&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG.&lt;br /&gt;
Deleted row&lt;br /&gt;
&lt;br /&gt;
== (#3) BRECORD_NOTES contains rows where BREC_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4 rows in BRECORD_NOTES where the BREC_FOL_date column, supposedly a date, contains time values that are not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_FOL_date&amp;quot;::TIME &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#4) MATING_EVENT contains rows where M_FOL_date has values that are not just a date ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The MATING_EVENT table contains 1 row where the time portion of M_FOL_date is not &amp;#039;00:00:00&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from &amp;quot;MATING_EVENT&amp;quot; where &amp;quot;M_FOL_date&amp;quot;::TIME WITHOUT TIME ZONE &amp;lt;&amp;gt; &amp;#039;00:00:00&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Fixed in MS Access, 11/16/2023, ICG&lt;br /&gt;
Deleted time stamp&lt;br /&gt;
&lt;br /&gt;
== (#5) BIOGRAPHY.DepartdateError data discarded ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The BIOGRAPHY.DepartdateError column contains data that was, after discussion, determined to be unusable.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_DepartdateError&amp;quot; &amp;lt;&amp;gt; 0;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Do not convert the data.  There is no corresponding column in the new db design.&lt;br /&gt;
&lt;br /&gt;
==  (#6) BIOGRAPHY.B_AnimID_num column contains the empty string ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains the empty string, instead of NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BIOGRAPHY&amp;quot; where &amp;quot;B_AnimID_num&amp;quot; = &amp;#039;&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed 11 empty string values to NULL in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#7) BIOGRAPHY.B_AnimID_num column is textual ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.B_AnimID_num column contains has a data type of TEXT.&lt;br /&gt;
The data values all begin with &amp;quot;CH&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_AnimID_num&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_AnimID_num&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        AND &amp;quot;B_AnimID_num&amp;quot; IS DISTINCT FROM NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Removed the &amp;quot;CH&amp;quot; prefix and made the column an integer in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#8) BIOGRAPHY.DadID_publication_info column contains NULL values ==&lt;br /&gt;
&lt;br /&gt;
No longer a problem as NULLs are now allowed.&lt;br /&gt;
&lt;br /&gt;
=== (Not) Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID_publication_info column has a data type that allows &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values.&lt;br /&gt;
SokweDB wants only text&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;This query no longer reports results because the data was changed in the original MS Access data.&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; IS NULL&lt;br /&gt;
  order by &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changed the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; to the empty string (&amp;quot;&amp;quot;) in MS Access. ICG 12/6/2023&lt;br /&gt;
&lt;br /&gt;
== (#9) BIOGRAPHY.DadID column contains non-AnimID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The BIOGRAPHY.DadID column has a non-AnimID values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;quot;B_AnimID&amp;quot;, &amp;quot;B_DadID&amp;quot;&lt;br /&gt;
  from raw.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
  where &amp;quot;B_DadID&amp;quot; is not null&lt;br /&gt;
        and &amp;quot;B_DadID&amp;quot; &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from raw.&amp;quot;BIOGRAPHY&amp;quot; as search&lt;br /&gt;
                          where search.&amp;quot;B_AnimID&amp;quot; = raw.&amp;quot;BIOGRAPHY&amp;quot;.&amp;quot;B_DadID&amp;quot;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create the DadIDPrelim column in a BIOGRAPHY_DATA table, and make a BIOGRAPHY view that combines Dad_ID and DadIDPrelim into DadID -- adding the &amp;#039;_prelim&amp;#039; suffix as expected.&lt;br /&gt;
&lt;br /&gt;
== (#10) BRECORD_NOTES contains rows where BREC_time has values that are not just a time ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 rows in BRECORD_NOTES where the BREC_time column, supposedly a time, contains date values that are not &amp;#039;1899-12-30&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from raw.&amp;quot;BRECORD_NOTES&amp;quot; where &amp;quot;BREC_time&amp;quot;::DATE &amp;lt;&amp;gt; &amp;#039;1899-12-30&amp;#039;;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
The data is fixed in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#11) KAZ has b_dadid_publication_info, but b_dadid is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has dad id publication info of &amp;#039;Rudicell et al. 2010&amp;#039;, but a &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; dadid.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * &lt;br /&gt;
  from clean.biography&lt;br /&gt;
  where b_dadid is NULL&lt;br /&gt;
        and (b_dadid_publication_info &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
             or b_dadid_publication_info is null);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; has 2 potential dads.  Introduce a DadIDStatus column, to replace the DadIDPrelim column and have a code that describes what&amp;#039;s going on with &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fixed in the MS Access data; &amp;lt;code&amp;gt;KAZ&amp;lt;/code&amp;gt; was assigned the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; individual as the dad.  Problem will be marked resolved with the upload of the next MS Access database dump.&lt;br /&gt;
&lt;br /&gt;
== (#12) 9 BIOGRAPHY rows have BirthComm values that are not COMM_IDS.CommID values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9 BIOGRAPHY rows with BirthComm values that are not valid COMM_IDS, their communities do not exist.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.biography&lt;br /&gt;
  where b_birthgroup is not NULL&lt;br /&gt;
        and not exists (select 1&lt;br /&gt;
                          from easy.community_lookup&lt;br /&gt;
                          where community_lookup.cl_community_id&lt;br /&gt;
                                = biography.b_birthgroup);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Added KL and KL_KK to CommIds.  KL = Kalande, KL_KK= Kasekela/kalande&lt;br /&gt;
codes are created during conversion&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 5d85b83bfe1a.&lt;br /&gt;
&lt;br /&gt;
== * (#13) COMM_MEMBS rows place individuals in a community, that is not their birth community, before their BIOGRAPHY.EntryDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
13 COMM_MEMBS rows place individuals into a community, that is not their birth community, before BIOGRAPHY says they&lt;br /&gt;
entered the community.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update: 12 rows as of 2026-07-19&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthgroup, b.b_entrydate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where b.b_birthgroup is distinct from cm.cm_cl_community_id&lt;br /&gt;
        and cm.cm_start_date &amp;lt; b.b_entrydate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Talk to Karl about kk_P0 and kk_P1&lt;br /&gt;
&lt;br /&gt;
== * (#14) COMM_MEMBS rows place individuals in a community after their BIOGRAPHY.EndDate ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
3 COMM_MEMBS rows place individuals into a community after BIOGRAPHY says they&lt;br /&gt;
left the community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, cm.cm_end_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_end_date &amp;gt; b.b_departdate&lt;br /&gt;
  order by b.b_animid, cm.cm_end_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG EVL fixed BH, HAI, TZB2 in Access. Need to talk to Karl about MG, RO, WD. Treat like KL chimps in Biography?&lt;br /&gt;
&lt;br /&gt;
== (#15) TT is placed in a community twice on the same day ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
An individual may not be in more than one community (or even twice in the same community) on any given day.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select first.cm_b_animid as anim_id&lt;br /&gt;
     , first.cm_start_date as first_start_date&lt;br /&gt;
     , first.cm_end_date as first_end_date&lt;br /&gt;
     , first.cm_cl_community_id as first_community_id&lt;br /&gt;
     , first.cm_start_source as first_start_source&lt;br /&gt;
     , first.cm_end_source as first_end_source&lt;br /&gt;
     , second.cm_start_date as second_start_date&lt;br /&gt;
     , second.cm_end_date as second_end_date&lt;br /&gt;
     , second.cm_cl_community_id as second_community_id&lt;br /&gt;
     , second.cm_start_source as second_start_source&lt;br /&gt;
     , second.cm_end_source as second_end_source&lt;br /&gt;
  from clean.community_membership as first&lt;br /&gt;
    join clean.community_membership as second&lt;br /&gt;
         on (first.cm_b_animid = second.cm_b_animid&lt;br /&gt;
             and first.cm_start_date &amp;lt; second.cm_start_date)&lt;br /&gt;
  where first.cm_end_date &amp;gt;= second.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access 2/1/2024 &lt;br /&gt;
Changed end date of KK_P1 membership to 8/14/2022&lt;br /&gt;
&lt;br /&gt;
ICG&lt;br /&gt;
&lt;br /&gt;
== (#16) There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow NULL values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
See problem #18.  Fixed in commit 7a48fb2c4aad3b5bd4&lt;br /&gt;
&lt;br /&gt;
== (#17) There are 64 rows in COMM_MEMB_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 rows in COMM_MEMB_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select date_of_update, chimp_id&lt;br /&gt;
  from clean.community_membership_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit 4c9626b304.&lt;br /&gt;
&lt;br /&gt;
== (#18) There are 31 rows in COMM_MEMB_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 31 rows in COMM_MEMB_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.community_membership_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
&lt;br /&gt;
This allows us to make the unknown person &amp;quot;inactive&amp;quot;, preventing them from being used&lt;br /&gt;
in newly entered data.  The alternative, allowing NULL MadeBy values, would allow&lt;br /&gt;
new &amp;quot;bad data&amp;quot;, and NULL values make querying harder.&lt;br /&gt;
&lt;br /&gt;
Fixed in commit  7a48fb2c4aad3b5b.&lt;br /&gt;
&lt;br /&gt;
== (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 18 rows in BIOGRAPHY_LOG where MadeBy is NULL, and the column does not allow&lt;br /&gt;
NULLs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where made_by is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create an unknown person (UNK), and when the person is NULL, use the unkonwn person.&lt;br /&gt;
(See Problem #18)== * (#19) There are 18 rows in BIOGRAPHY_LOG where the MadeBy is NULL ==&lt;br /&gt;
&lt;br /&gt;
Fixed in commit 8a528cc7d1a9fb.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update 2026-07-26:&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
The 18 problematic values are zero-length strings, not NULL, so they pass through the ELSE made_by branch and violate the foreign key as madeby=(). This solution was hardened with commit 6279986f5f77cab69d19733925f173eadca23b60.&lt;br /&gt;
&lt;br /&gt;
== (#20) There are 3 rows in BIOGRAPHY_LOG where the Rationale is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 rows in BIOGRAPHY_LOG where Rationale is NULL.  The column does not allow&lt;br /&gt;
NULLs; normally the conversion program would convert NULL to the empty string.&lt;br /&gt;
But the Rationale column requires there be (non-empty) textual data.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_rationale is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024. Updated rationale to &amp;#039;routine update&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
== (#21) There are 6 rows in BIOGRAPHY_LOG where the update_escription is NULL ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 6 rows in BIOGRAPHY_LOG where Description is NULL.&lt;br /&gt;
It appears that the description was put into the Rationale column, sometimes along with some rationale.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.biography_update_log where update_description is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
ICG fixed in MS Access 2/8/2024, using information in update_rationale&lt;br /&gt;
&lt;br /&gt;
== (#22) There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not on BIOGRAPHY ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 85 rows in BIOGRAPHY_LOG where the chimp_id is not a BIOGRAPHY.AnimID.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.biography&lt;br /&gt;
                      where biography.b_animid = log.chimp_id);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Make this a &amp;quot;soft&amp;quot; error.  Fixed in commit c96555f9f326.&lt;br /&gt;
&lt;br /&gt;
== (#23) There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 65 rows in BIOGRAPHY_LOG where the MadeBy is not on PEOPLE, but there is MadeBy data.&lt;br /&gt;
&lt;br /&gt;
These are &amp;quot;combination&amp;quot; ID errors, where multiple people are entered instead&lt;br /&gt;
of a single people code.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
In the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema run:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.biography_update_log as log&lt;br /&gt;
  where not exists (select 1&lt;br /&gt;
                      from clean.people&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
                      where people.person = log.made_by)&lt;br /&gt;
        and made_by is not null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Fixed in MS Access by ICG 2/8/2024. Changed multiple IDs to the one who made the final change.&lt;br /&gt;
&lt;br /&gt;
== (#24) Follow starts are not first arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,397 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_begin&amp;quot; is not the first arrival time, fa_time_start,&lt;br /&gt;
on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow start is supposed&lt;br /&gt;
to be the first arrival.  Is there additional data&lt;br /&gt;
in the old fol_time_begin, like the actual start&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.min_start = follow.fol_time_begin)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_begin.&lt;br /&gt;
&lt;br /&gt;
==(#25) Follow ends are not last arrivals ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 6,570 cases where the old FOLLOW table&amp;#039;s columns &lt;br /&gt;
&amp;quot;fol_time_end&amp;quot; is not the last arrival&lt;br /&gt;
time, fa_time_end, on FOLLOW_ARRIVAL.&lt;br /&gt;
&lt;br /&gt;
The The Gombe Chimpanzee Database Handbook says that this&lt;br /&gt;
should never happen, because the follow end is supposed&lt;br /&gt;
to be the first arrival/last departure.  Is there additional data&lt;br /&gt;
in the old fol_time_end, like the actual end-time&lt;br /&gt;
the observers started working?  If not, which data is correct,&lt;br /&gt;
the one in FOLLOW or the one in FOLLOW_ARRIVAL?&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
          , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM clean.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_end&lt;br /&gt;
  FROM spans&lt;br /&gt;
    JOIN clean.follow&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
        (SELECT 1&lt;br /&gt;
           FROM clean.follow&lt;br /&gt;
           WHERE follow.fol_date = spans.fa_fol_date&lt;br /&gt;
                 AND follow.fol_b_animid = spans.fa_fol_b_focal_animid&lt;br /&gt;
                 AND spans.max_end = follow.fol_time_end)&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is data entry error, ignore the problem and ignore the data in the follow.fol_time_end.&lt;br /&gt;
&lt;br /&gt;
== (#26) Mismatch of start-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;1,848&amp;lt;/s&amp;gt; 3,369 follows where the follow arrival says the focal started in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MIN(fa_time_start) AS min_start&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*, follow.fol_time_begin, follow.fol_flag_begin_in_nest&lt;br /&gt;
     , first_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS first_arrivals&lt;br /&gt;
      ON (first_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND first_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND first_arrivals.fa_time_start = spans.min_start)&lt;br /&gt;
  WHERE (((first_arrivals.fa_type_of_nesting = 1&lt;br /&gt;
           OR first_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_begin_in_nest = 0)&lt;br /&gt;
         OR (first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 1&lt;br /&gt;
             AND first_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_begin_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore. Default to data from Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
When applying the solution, solving both problem #26 and #27,&lt;br /&gt;
and updating follow_arrival so that it is the one source of truth&lt;br /&gt;
used by the conversion,&lt;br /&gt;
there are 297 follow_arrival rows updated.&lt;br /&gt;
&lt;br /&gt;
This implies that it is primarily the follow table&lt;br /&gt;
that does not mark the individual as being in a&lt;br /&gt;
nest when they should be in a nest.  (&amp;quot;Should be&amp;quot;,&lt;br /&gt;
according to the accepted solution.)&lt;br /&gt;
&lt;br /&gt;
==(#27) Mismatch of end-in-nest on follow and follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3,804 follows where the follow arrival says the focal ended in the nest but the follow does not, or vice-versa.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH spans AS&lt;br /&gt;
  (SELECT fa_fol_date, fa_fol_b_focal_animid&lt;br /&gt;
        , MAX(fa_time_end) AS max_end&lt;br /&gt;
     FROM easy.follow_arrival&lt;br /&gt;
     WHERE fa_b_arr_animid = fa_fol_b_focal_animid&lt;br /&gt;
     GROUP BY fa_fol_date, fa_fol_b_focal_animid)&lt;br /&gt;
SELECT spans.*&lt;br /&gt;
     , follow.fol_time_end, follow.fol_flag_end_in_nest&lt;br /&gt;
     , last_arrivals.fa_type_of_nesting&lt;br /&gt;
  FROM easy.follow&lt;br /&gt;
    JOIN spans&lt;br /&gt;
      ON (follow.fol_date = spans.fa_fol_date&lt;br /&gt;
          AND follow.fol_b_animid = spans.fa_fol_b_focal_animid)&lt;br /&gt;
    JOIN easy.follow_arrival&lt;br /&gt;
      AS last_arrivals&lt;br /&gt;
      ON (last_arrivals.fa_fol_date = follow.fol_date&lt;br /&gt;
          AND last_arrivals.fa_fol_b_focal_animid = follow.fol_b_animid&lt;br /&gt;
          AND last_arrivals.fa_time_end = spans.max_end)&lt;br /&gt;
  WHERE (((last_arrivals.fa_type_of_nesting = 2&lt;br /&gt;
           OR last_arrivals.fa_type_of_nesting = 3)&lt;br /&gt;
          AND follow.fol_flag_end_in_nest = 0)&lt;br /&gt;
         OR (last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 2&lt;br /&gt;
             AND last_arrivals.fa_type_of_nesting &amp;lt;&amp;gt; 3&lt;br /&gt;
             AND follow.fol_flag_end_in_nest = 1))&lt;br /&gt;
  ORDER BY spans.fa_fol_date, spans.fa_fol_b_focal_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem. Default to Follow_Arrival&lt;br /&gt;
&lt;br /&gt;
==== Remarks ====&lt;br /&gt;
&lt;br /&gt;
See the remarks for [[Conversion Data Issues#(#27) Mismatch of end-in-nest on follow and follow_arrival|problem #26]].&lt;br /&gt;
&lt;br /&gt;
== (#28) The FOLLOW.FOL_distance_traveled column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.FOL_distance_traveled column has no corresponding column in the new&lt;br /&gt;
database design.  The conversion process does not check that the value&lt;br /&gt;
of this column is consistent with the other data in the database from which it is computed.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This is a computed column and does not need to be converted.  The desired value is computed in the design of the new database.&lt;br /&gt;
&lt;br /&gt;
The assumption is that the &amp;quot;raw&amp;quot; data from which this value is computed in the MS Access database is correct.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#29) The FOLLOW.Brecord_notes column is not converted ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
The FOLLOW.Brecord_notes column has no corresponding column in the new&lt;br /&gt;
database design.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
This column is used for administrative purposes and does not need to be in the new database design.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#30) Some follows have no community ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 13 follows with no community.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where fol_cl_community_id is null&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#31) Some follows have an animid with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 39 follows with animids that don&amp;#039;t exist, because they have trailing spaces.  These are comprised of &amp;lt;s&amp;gt;9&amp;lt;/s&amp;gt; 23 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select rtrim(fol_b_animid)&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where rtrim(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in the conversion process.&lt;br /&gt;
&lt;br /&gt;
== (#32) Some follows have an animid that is lower-case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follows with an animids that don&amp;#039;t exist, because it is lower-case.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned up in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from easy.follow&lt;br /&gt;
  where upper(fol_b_animid) &amp;lt;&amp;gt; fol_b_animid&lt;br /&gt;
  order by fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Convert the animid to upper-case in the conversion process.&lt;br /&gt;
&lt;br /&gt;
==(#33) Some follows have an animid that does not exist, even after animid cleanup ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14 follows with animids that don&amp;#039;t exist, even after cleanup that removes trailing spaces and forces upper-case.  (Some of these may be due to prior errors.)  These are comprised of 3 distinct animids.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 14 follows with bad animids&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The 3 animids involved&lt;br /&gt;
select distinct follow.fol_b_animid&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where not exists&lt;br /&gt;
          (select 1&lt;br /&gt;
             from clean.biography&lt;br /&gt;
             where biography.b_animid = upper(rtrim(follow.fol_b_animid)))&lt;br /&gt;
  order by follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#34) There are duplicate animid, date combinations on FOLLOW ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2 sets of duplicate animid, date combinations on the FOLLOW table, for a total of 4 rows.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate animid, date combinations&lt;br /&gt;
select follow.fol_b_animid, follow.fol_date, count(*)&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
  having count(*) &amp;gt; 1&lt;br /&gt;
  order by fol_b_animid, fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The duplicate rows&lt;br /&gt;
with dups as (&lt;br /&gt;
  select follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    from clean.follow&lt;br /&gt;
    group by follow.fol_b_animid, follow.fol_date&lt;br /&gt;
    having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
       join dups on (follow.fol_b_animid = dups.fol_b_animid&lt;br /&gt;
                     and follow.fol_date = dups.fol_date)&lt;br /&gt;
  order by follow.fol_b_animid, follow.fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#35) There are follows done before a focal was under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5 follows that are done before their focal was under study, before the focal&amp;#039;s EntryDate.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow.*, biography.b_entrydate&lt;br /&gt;
  from clean.follow&lt;br /&gt;
    join clean.biography on (biography.b_animid = follow.fol_b_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow.fol_date&lt;br /&gt;
  order by follow.fol_date, follow.fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#36) FOLLOW_OBSERVERS.Period is not checked against follow start or stop times ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The conversion process uses the clean.follow columns of fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, and fol_pm_observer_2 to populate FOLLOW_OBSERVERS.&lt;br /&gt;
(The *_observer_1 column going into FOLLOW_OBSERVERS.OBS_BRec and the *_observer_2 column going in OBS_Tiki.)&lt;br /&gt;
&lt;br /&gt;
If both the *_observer_1 and the *_observer_2 columns are either NULL or the empty string (after space trimming), the no row is created for the respective time period.&lt;br /&gt;
&lt;br /&gt;
There are no checks done to ensure that the time periods of the follow have any relation to the FOLLOW_OBSERVERS.Period value.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  (If someone cares, add a query to the warning system.)&lt;br /&gt;
&lt;br /&gt;
== (#37) Some follows have no recorded observers ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 56 follows with no recorded observers, but the system requires there be a FOLLOW_OBSERVERS record related to the follow.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow&lt;br /&gt;
  where coalesce(btrim(fol_am_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_am_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_1), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
        and  coalesce(btrim(fol_pm_observer_2), &amp;#039;&amp;#039;) = &amp;#039;&amp;#039;&lt;br /&gt;
  order by fol_date, fol_b_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make the &amp;quot;NONE&amp;quot; (no observer) person the observers.&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;UNK&amp;quot; (unknown) time period in PERIODS, and make that the time period.&lt;br /&gt;
&lt;br /&gt;
== (#38) Some follows have only one observer, but the system wants two ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are follows with only one observer, but the system requires a FOLLOW_OBSERVERS record have a value in both OBS_BRec and OBS_Tiki.&lt;br /&gt;
&lt;br /&gt;
See problem #36 for a description of the follow observer conversion process.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Create a &amp;quot;NONE&amp;quot; person, and make that person the observer when there&amp;#039;s otherwise not a value.&lt;br /&gt;
&lt;br /&gt;
== (#39) In follow, there are observers that are 2 people ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 30 rows (28, really, because of n/a) that appear to represent 2 people.&lt;br /&gt;
&lt;br /&gt;
These look like names, separated by the &amp;quot;/&amp;quot; character.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM clean.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
          AND STRPOS(uniq_people.person, &amp;#039;/&amp;#039;) &amp;lt;&amp;gt; 0&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ignore the problem.  Mark all people containing a &amp;quot;/&amp;quot; character as not active, so they can&amp;#039;t be used in the future.&lt;br /&gt;
&lt;br /&gt;
== (#40) In follow, there are observers that differ only by character case ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow table, in columns fol_am_observer_1, fol_am_observer_2, fol_pm_observer_1, fol_pm_observer_2, there are 511 names that differ only by character case.  So, about half that in terms of unique names.&lt;br /&gt;
&lt;br /&gt;
This is complicated to account for and exclude duplicates in the conversion process.  So resolution of this is holding back additional conversion work.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note the use of the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, instead of the&lt;br /&gt;
&amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.  This is because the observer data&lt;br /&gt;
in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema has been case-normalized.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH new_people AS (&lt;br /&gt;
  SELECT BTRIM(follow.fol_am_observer_1) AS person&lt;br /&gt;
    FROM easy.follow&lt;br /&gt;
    WHERE follow.fol_am_observer_1 IS NOT NULL&lt;br /&gt;
    GROUP BY BTRIM(follow.fol_am_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_am_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_am_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_am_observer_2)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_1) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_1 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_1)&lt;br /&gt;
  UNION&lt;br /&gt;
    SELECT BTRIM(follow.fol_pm_observer_2) AS person&lt;br /&gt;
      FROM easy.follow&lt;br /&gt;
      WHERE follow.fol_pm_observer_2 IS NOT NULL&lt;br /&gt;
      GROUP BY BTRIM(follow.fol_pm_observer_2)&lt;br /&gt;
)&lt;br /&gt;
, uniq_people AS (&lt;br /&gt;
  SELECT new_people.person&lt;br /&gt;
    FROM new_people&lt;br /&gt;
    WHERE new_people.person &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
    GROUP BY new_people.person&lt;br /&gt;
  )&lt;br /&gt;
  SELECT uniq_people.person&lt;br /&gt;
    FROM uniq_people&lt;br /&gt;
         JOIN uniq_people AS up&lt;br /&gt;
              ON (LOWER(uniq_people.person) = LOWER(up.person))&lt;br /&gt;
    WHERE uniq_people.person &amp;lt;&amp;gt; up.person&lt;br /&gt;
    GROUP BY uniq_people.person&lt;br /&gt;
    ORDER BY LOWER(uniq_people.person), uniq_people.person;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Use the mixed case code when such exists.  This involves changing the observer columns in the &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt; table.  See [[The_Old_Database#The_observer_columns_of_the_follow_table|the notes]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==(#41)There are follow arrivals with NULL nesting information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 11 rows with a NULL fa_type_of_nesting.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_nesting IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Solved&lt;br /&gt;
&lt;br /&gt;
== (#42) There are follow arrivals with NULL sexual cycle information ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are 222 rows with a NULL fa_type_of_cycle.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from clean.follow_arrival where fa_type_of_cycle is null;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Make a &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; &amp;lt;code&amp;gt;CYCLE_STATES&amp;lt;/code&amp;gt; value and use that for the arrivals with no cycle information.&lt;br /&gt;
&lt;br /&gt;
== * (#43) There are follow arrivals with no related follow ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;732&amp;lt;/s&amp;gt; 276 rows with a focal and a date that have no matching information on the follow table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
    (SELECT 1&lt;br /&gt;
       FROM clean.follow&lt;br /&gt;
       WHERE follow.fol_b_animid = follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
             AND follow.fol_date = follow_arrival.fa_fol_date)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
5/5/2026 ICG&lt;br /&gt;
FIXED MANY ROWS IN MS ACCESS&lt;br /&gt;
&lt;br /&gt;
Most of the rest are hand-entered juvenile arrivals from Brec. Spot checks have Brec notes but no physical tiki&lt;br /&gt;
Need to investigate further - brec swahili? Eventually the rest of the individuals in the Brec notes need to be entered by hand.&lt;br /&gt;
&lt;br /&gt;
Should be solved by the WATCHES table&lt;br /&gt;
&lt;br /&gt;
For now, allow these rows&lt;br /&gt;
&lt;br /&gt;
==== Remarks on Solution ====&lt;br /&gt;
&lt;br /&gt;
There is no follow to attach the arrival to.  No WATCHES row.&lt;br /&gt;
So there is no way to get the data in and simply write&lt;br /&gt;
a warning.&lt;br /&gt;
&lt;br /&gt;
Allowing these rows requires a re-design of the database&lt;br /&gt;
and additional &amp;quot;hard&amp;quot; rules.  Which is going to take time&lt;br /&gt;
to program, and delay finishing.  And, it will take additional time&lt;br /&gt;
to rip-out the rules when you finally resolve the data issue.&lt;br /&gt;
&lt;br /&gt;
What is needed is to make a new WATCHES.Type just to allow for arrivals&lt;br /&gt;
with no related follow, and ensure that only arrivals without&lt;br /&gt;
follows are attaching to this new Type.&lt;br /&gt;
&lt;br /&gt;
The only &amp;quot;easy&amp;quot; alternative I can think of is:&lt;br /&gt;
&lt;br /&gt;
  Don&amp;#039;t convert this data at this time.  Convert it when&lt;br /&gt;
  we get around to fixing the data.  In the mean time,&lt;br /&gt;
  anyone who wants to work with this data can get it&lt;br /&gt;
  from the &amp;quot;clean&amp;quot; schema, using the query provided in&lt;br /&gt;
  the problem #43 section, and figure out how to integrate&lt;br /&gt;
  it with whatever they are working on.&lt;br /&gt;
&lt;br /&gt;
==(#44) There are follow arrivals where the arriving chimp arrives before being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;217&amp;lt;/s&amp;gt; 49 rows where the fa_b_arr_animid, the arriving individual, has an entry date after the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_entrydate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_entrydate &amp;gt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Something is wrong with community entry date or biography&lt;br /&gt;
&lt;br /&gt;
SN1 fixed in MS Access ICG 5/19/2026&lt;br /&gt;
&lt;br /&gt;
Elo Confirmed CT and SG were with CA on 10/23/1980 from brec swahili&lt;br /&gt;
&lt;br /&gt;
RESOLVED&lt;br /&gt;
&lt;br /&gt;
== (#45) The follow_arrival.fa_update column is not preserved ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is nowhere in the current design to store the values in the follow_arrival.fa_update column.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Add a temporal extension to Postgres to make the db into a temporal database to track change history and be able to see data as it existed at any point in time.&lt;br /&gt;
&lt;br /&gt;
FOLLOWUP WITH KARL. &lt;br /&gt;
Will re-review this should a temporal extension be out of budget, etc.&lt;br /&gt;
&lt;br /&gt;
ASK KARL WHY THIS COLUMN CAN&amp;#039;T GO IN&lt;br /&gt;
&lt;br /&gt;
There is now a ARRIVALS.Updated column.  This provides a place to put the existing&lt;br /&gt;
data.  Going forward, a PG extension (see below) can be installed for more functionality.&lt;br /&gt;
&lt;br /&gt;
Wouldn&amp;#039;t you rather have a [https://en.wikipedia.org/wiki/Temporal_database temporal database]?  (Simple description in the top paragraph [https://wiki.postgresql.org/wiki/Temporal_Extensions here].)  This would allow you to &amp;quot;time travel&amp;quot; and look at the database content at any point in time.  And, it puts a &amp;quot;last changed&amp;quot; timestamp on every row, of every table you say to track temporally, automatically.&lt;br /&gt;
&lt;br /&gt;
If so, install one of [https://wiki.postgresql.org/wiki/Temporal_Extensions these extensions] to PostgreSQL.  (Or, as a separate project, I will improve one of them so that the people who don&amp;#039;t need to time travel don&amp;#039;t see any changes to the existing database at all.  What I don&amp;#039;t like about most of the extensions listed is that they add extra columns to existing tables, as well as adding extra &amp;quot;history&amp;quot; tables in places where they clutter up what users really need to see.)&lt;br /&gt;
&lt;br /&gt;
== (#46) There are follow_arrivals where non-females have a cycle code that is other than &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;169&amp;lt;/s&amp;gt; 168 follow_arrival rows, for non-female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) that is not &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To summarize by sex and cycle code:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select biography.b_sex, follow_arrival.fa_type_of_cycle, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex &amp;lt;&amp;gt; &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;n/a&amp;#039;&lt;br /&gt;
  group by biography.b_sex, follow_arrival.fa_type_of_cycle order by biography.b_sex&lt;br /&gt;
         , follow_arrival.fa_type_of_cycle;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
FIX THE FEW ACTUAL MALES THAT DON&amp;#039;T HAVE AN n/a&lt;br /&gt;
&lt;br /&gt;
Change non-females (INCLUDING UNKNOWN SEX) who have a cycle code of &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; to a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;. IAN TO FIX IN ACCESS&lt;br /&gt;
&lt;br /&gt;
IAN FIXED IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== * (#47) There are follow_arrivals where females that are too young have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 16 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are less than 5 years of age.&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;5 years&amp;#039;::INTERVAL&lt;br /&gt;
                )&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note: The original query used the wrong inequality symbol.&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Changing the limit from 6 years of age to 5 reduced the number of outstanding problems to 16.&lt;br /&gt;
&lt;br /&gt;
The rest will have to be fixed in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 5 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
Resolution: Fixed in the data.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23 it does not appear that these data were fixed, please see below for a summary of this issue and related Problem #49:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== * (#48) There are follow arrivals where the arriving chimp arrives after finishing being under study ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
In the follow_arrival table, there are &amp;lt;s&amp;gt;239&amp;lt;/s&amp;gt; 324 rows where the fa_b_arr_animid, the arriving individual, has a departure date before the date of the follow.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_departdate, biography.b_sex&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
         on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_departdate &amp;lt; follow_arrival.fa_fol_date&lt;br /&gt;
  order by biography.b_animid, follow_arrival.fa_fol_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN TO FIX IN ACCESS - CHECK DATES OF FOLLOWS AND ACCURACY OF IDS&lt;br /&gt;
&lt;br /&gt;
== * (#49) There are follow_arrivals where females that are too old have a cycle code of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;127&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;98&amp;lt;/s&amp;gt; 86 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt; but are more than 14 years of age.&lt;br /&gt;
(Actually, because the endpoint takes up the whole 14th year, this means at least 15 years of age.)&lt;br /&gt;
&lt;br /&gt;
Naturally, this number will change if the age limit is changed, but this is here as a placeholder.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the maximum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Note: Original query used the wrong inequality.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_type_of_cycle = &amp;#039;U&amp;#039;&lt;br /&gt;
	and biography.b_birthdate&lt;br /&gt;
              &amp;lt;= (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;1 year&amp;#039;::INTERVAL&lt;br /&gt;
                 - &amp;#039;14 years&amp;#039;::INTERVAL&lt;br /&gt;
		)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
The number of problem rows was reduced after changing the limit to 14 years from 9 years.&lt;br /&gt;
&lt;br /&gt;
The remaining problems will have to be adjusted in the data.  Alternately, we can adjust the hard limit, and set a soft limit of 14 years in the warning system with a note to change the hard limit back once the errors are resolved.&lt;br /&gt;
&lt;br /&gt;
5/7/2026: ICG checked through 2003. When there was a &amp;quot;?&amp;quot; for swelling, I entered &amp;#039;0&amp;#039;. If there was actually a &amp;#039;U&amp;#039; on the tiki itself, I left &amp;quot;U&amp;quot; in the FA table. Note that a lot of these females are MGF, so the U might be legit.&lt;br /&gt;
&lt;br /&gt;
5/19/2026: MGFs have a dummy birthdate which is why the age is being flagged. The rest are what the observer recorded, so let them in.&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-23: There are 86 failing rows; many but not all are MFG (how to we treat MFG?). Please see below for a summary of this issue and related Problem #47:&lt;br /&gt;
&amp;#039;&amp;#039;note that this summary does not include MFG&amp;#039;&amp;#039;&lt;br /&gt;
{| class=&amp;quot;wikitable sortable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! problem !! problem_description !! age_years !! current_cycle !! youngest_exact_age !! oldest_exact_age !! offending_rows !! problem_total&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 1 || U || 1 year 4 mons 5 days || 1 year 4 mons 5 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 2 || U || 2 years 11 mons 26 days || 2 years 11 mons 26 days || 1 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 3 || U || 3 years 3 mons 30 days || 3 years 11 mons 1 day || 10 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 47 || U assigned before age 5 || 4 || U || 4 years 5 mons 21 days || 4 years 9 mons 18 days || 9 || 21&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 15 || U || 15 years 1 mon 2 days || 15 years 10 mons 14 days || 16 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 16 || U || 16 years 19 days || 16 years 11 mons 6 days || 7 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 17 || U || 17 years 1 mon 19 days || 17 years 1 mon 19 days || 1 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 18 || U || 18 years 1 mon 18 days || 18 years 8 mons 16 days || 6 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 32 || U || 32 years 1 mon 21 days || 32 years 5 mons 14 days || 4 || 38&lt;br /&gt;
|-&lt;br /&gt;
| 49 || U assigned at age 15 or older || 33 || U || 33 years 6 days || 33 years 6 days || 4 || 38&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== (#50) There are follow_arrivals where females have a cycle code of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,496 follow_arrival rows, for female arriving individuals, that have a sexual cycle code (fa_type_of_cycle) of &amp;lt;code&amp;gt;n/a&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
     , biography.b_sex&lt;br /&gt;
     , biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle = &amp;#039;n/a&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change the cycle code to &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt;, for these rows.  This is done in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
This is an indication that future cleanup is required.&lt;br /&gt;
&lt;br /&gt;
== (#51) There are follow_arrivals with invalid fa_data_source values  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 546 follow_arrival rows that have fa_data_source values that are not one of:&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_Mom&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Tiki_ID&amp;lt;/code&amp;gt;&lt;br /&gt;
&amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Query the &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema, because the data has been cleaned in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&lt;br /&gt;
-- Summarize with:&lt;br /&gt;
select follow_arrival.fa_data_source, count(*)&lt;br /&gt;
  from easy.follow_arrival&lt;br /&gt;
  where follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_Mom&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Tiki_ID&amp;#039;&lt;br /&gt;
	and follow_arrival.fa_data_source &amp;lt;&amp;gt; &amp;#039;Brec&amp;#039;&lt;br /&gt;
  group by follow_arrival.fa_data_source&lt;br /&gt;
  order by follow_arrival.fa_data_source;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;BREC&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;brec&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Brec&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Change &amp;lt;code&amp;gt;TikI&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;Tiki&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Add the other codes:&lt;br /&gt;
&lt;br /&gt;
  fa_data_source	count&lt;br /&gt;
  Tiki_GM	202&lt;br /&gt;
  Tiki_PM	165&lt;br /&gt;
  Tiki_SS	22&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also: Tiki_Mom&lt;br /&gt;
&lt;br /&gt;
== * (#52) There are follow_arrivals that are almost duplicates  ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
This entry may end up being multiple problems.&lt;br /&gt;
&lt;br /&gt;
There are follow_arrivals that are near duplicates.&lt;br /&gt;
When checking for duplicates on &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_seq_num&amp;lt;/code&amp;gt;, the rows are always unique.&lt;br /&gt;
But checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt;,&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; yields 184 rows.&lt;br /&gt;
Leaving off &amp;lt;code&amp;gt;fa_time_end&amp;lt;/code&amp;gt; and just checking the combination of &amp;lt;code&amp;gt;fa_fol_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_fol_b_focal_animid&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt;&lt;br /&gt;
and &amp;lt;code&amp;gt;fa_time_start&amp;lt;/code&amp;gt; yields 430 rows.&lt;br /&gt;
&lt;br /&gt;
Why the duplicates?&lt;br /&gt;
&lt;br /&gt;
This entry is a call for a definition of an &amp;lt;code&amp;gt;ARRIVALS&amp;lt;/code&amp;gt; row, what does it mean to be a duplicate?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- no duplicates when looking at just sequence number&lt;br /&gt;
select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_seq_num&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_seq_num&lt;br /&gt;
     having count(*) &amp;gt; 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Checking against start and end time&lt;br /&gt;
with dups as&lt;br /&gt;
  (select fa_fol_date&lt;br /&gt;
        , fa_fol_b_focal_animid&lt;br /&gt;
        , fa_b_arr_animid&lt;br /&gt;
        , fa_time_start&lt;br /&gt;
        , fa_time_end&lt;br /&gt;
     from clean.follow_arrival&lt;br /&gt;
     group by fa_fol_date&lt;br /&gt;
            , fa_fol_b_focal_animid&lt;br /&gt;
            , fa_b_arr_animid&lt;br /&gt;
            , fa_time_start&lt;br /&gt;
            , fa_time_end&lt;br /&gt;
     having count(*) &amp;gt; 1)&lt;br /&gt;
select *&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where exists&lt;br /&gt;
    (select 1                                             &lt;br /&gt;
       from dups                                     &lt;br /&gt;
       where follow_arrival.fa_fol_date = dups.fa_fol_date&lt;br /&gt;
             and follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
                 = dups.fa_fol_b_focal_animid&lt;br /&gt;
             and follow_arrival.fa_b_arr_animid = dups.fa_b_arr_animid&lt;br /&gt;
             and follow_arrival.fa_time_start = dups.fa_time_start&lt;br /&gt;
             and follow_arrival.fa_time_end = dups.fa_time_end)&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end&lt;br /&gt;
         , follow_arrival.fa_seq_num;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
Where both start and end-times are duplicated, along with other data like cycle state, I can add additional observers or data sources.  But I don&amp;#039;t know what to do with &amp;quot;almost the same&amp;quot; data.&lt;br /&gt;
&lt;br /&gt;
FOR NOW, PROMPT A WARNING.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
I leave it to all of you to decide just what the warning(s) are to report.&lt;br /&gt;
&lt;br /&gt;
== * (#53) There are follow_arrivals where the arriving chimp does not exist ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;688&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;84&amp;lt;/s&amp;gt; 78 rows, having &amp;lt;s&amp;gt;3&amp;lt;/s&amp;gt; &amp;lt;s&amp;gt;6&amp;lt;/s&amp;gt; 2 different &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; values, where the &amp;lt;code&amp;gt;fa_b_arr_animid&amp;lt;/code&amp;gt; value is not a &amp;lt;code&amp;gt;biography.b_animid&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- The rows&lt;br /&gt;
select *                                       &lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid;&lt;br /&gt;
&lt;br /&gt;
-- A summary of the bad animal ids&lt;br /&gt;
select follow_arrival.fa_b_arr_animid, count(*)&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
  where not exists&lt;br /&gt;
    (select 1&lt;br /&gt;
       from clean.biography&lt;br /&gt;
       where biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  group by follow_arrival.fa_b_arr_animid&lt;br /&gt;
  order by follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in MS Access 10/2025. AMA--&amp;gt;AME, OBE--&amp;gt;POR&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
SRE: still need to address: UWE, GGl, FAC, FAR, GU, and Sl (as of the April 2026 dump)&lt;br /&gt;
&lt;br /&gt;
FIXED UWE, GGl, FAC, FAR, GU, and Sl IN ACCESS 7/3/2026&lt;br /&gt;
&lt;br /&gt;
UPDATE 2026-07-22: still need to resolve FAR (17 records) and FAC (61 records)&lt;br /&gt;
&lt;br /&gt;
== * (#54) There are follow_arrivals where females that are too young have a cycle state code that is not &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;U&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;MISS&amp;lt;/code&amp;gt; ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;s&amp;gt;26&amp;lt;/s&amp;gt; 42 follow_arrival rows, for female arriving individuals, that have a sexual cycle code indicating sexual swelling but are less than 8 years of age.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select follow_arrival.*, biography.b_sex, biography.b_birthdate&lt;br /&gt;
  from clean.follow_arrival&lt;br /&gt;
    join clean.biography&lt;br /&gt;
           on (biography.b_animid = follow_arrival.fa_b_arr_animid)&lt;br /&gt;
  where biography.b_sex = &amp;#039;F&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;0&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;U&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_type_of_cycle &amp;lt;&amp;gt; &amp;#039;MISS&amp;#039;&lt;br /&gt;
        and biography.b_birthdate&lt;br /&gt;
              &amp;gt; (follow_arrival.fa_fol_date&lt;br /&gt;
                 - &amp;#039;8 years&amp;#039;::INTERVAL&lt;br /&gt;
                 )&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF2&amp;#039;&lt;br /&gt;
        and follow_arrival.fa_b_arr_animid &amp;lt;&amp;gt; &amp;#039;MGF3&amp;#039;&lt;br /&gt;
  order by follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN TO FIX - CHANGE ALL TO &amp;quot;U&amp;quot;&lt;br /&gt;
&lt;br /&gt;
That works.  Would it be better to leave the original MS Access data untouched and have the conversion process make the change in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
==(#55) There are community_membership rows that place an individual in a community before birth ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one individual who is placed in a community, once, before birth.&lt;br /&gt;
&lt;br /&gt;
Note: The test is against the birthdate, not the minimum possible birthdate.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select b.b_animid, b.b_birthdate, cm.cm_start_date&lt;br /&gt;
  from clean.community_membership as cm&lt;br /&gt;
    join clean.biography as b&lt;br /&gt;
         on (b.b_animid = cm.cm_b_animid)&lt;br /&gt;
  where cm.cm_start_date &amp;lt; b.b_birthdate&lt;br /&gt;
  order by b.b_animid, cm.cm_start_date;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Bad solution ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed FN community starte date in MS Access - oct 2025&lt;br /&gt;
&lt;br /&gt;
== (#56) There are follow_arrival focal animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 144 follow arrivals where the focal id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from follow_arrival where fa_fol_b_focal_animid &amp;lt;&amp;gt; rtrim(fa_fol_b_focal_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#57) GROOM_BOUT duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;GROOM_BOUT&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;GRM_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_FOL_B_focal_AnimId&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_time_begin&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;GRM_B_partner_AnimId&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;GRM_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
               , &amp;quot;GRM_time_begin&amp;quot; AS the_time&lt;br /&gt;
               , &amp;quot;GRM_B_partner_AnimId&amp;quot; AS the_partner&lt;br /&gt;
            FROM raw.&amp;quot;GROOM_BOUT&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;GRM_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_time_begin&amp;quot;&lt;br /&gt;
                   , &amp;quot;GRM_B_partner_AnimId&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS gb&lt;br /&gt;
      ON (&amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_date&amp;quot; = gb.the_date&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_FOL_B_focal_AnimId&amp;quot; = gb.the_animid&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_time_begin&amp;quot; = gb.the_time&lt;br /&gt;
          AND &amp;quot;GROOM_BOUT&amp;quot;.&amp;quot;GRM_B_partner_AnimId&amp;quot; = gb.the_partner);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG TO LOOK AT SOURCE OF PROBLEM. MAYBE JUST DELETE DUPLICATES?&lt;br /&gt;
&lt;br /&gt;
== * (#58) OTHER_SPECIES duplicate keys ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
The data dump says that the &amp;lt;code&amp;gt;OTHER_SPECIES&amp;lt;/code&amp;gt; table has a primary key consisting of, in order, the columns: &amp;lt;code&amp;gt;OS_FOL_date&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_FOL_B_focal_AnimID&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;OS_time_begin&amp;lt;/code&amp;gt;&lt;br /&gt;
But these columns contain duplicate values.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The duplicate values can be listed (from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
    JOIN (SELECT &amp;quot;OS_FOL_date&amp;quot; AS the_date&lt;br /&gt;
               , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; AS the_animid&lt;br /&gt;
	       , &amp;quot;OS_time_begin&amp;quot; AS the_time&lt;br /&gt;
            FROM raw.&amp;quot;OTHER_SPECIES&amp;quot;&lt;br /&gt;
            GROUP BY &amp;quot;OS_FOL_date&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_FOL_B_focal_AnimId&amp;quot;&lt;br /&gt;
                   , &amp;quot;OS_time_begin&amp;quot;&lt;br /&gt;
            HAVING count(*) &amp;gt; 1&lt;br /&gt;
         ) AS os&lt;br /&gt;
      ON (&amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_date&amp;quot; = os.the_date&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_FOL_B_focal_AnimId&amp;quot; = os.the_animid&lt;br /&gt;
          AND &amp;quot;OTHER_SPECIES&amp;quot;.&amp;quot;OS_time_begin&amp;quot; = os.the_time);&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Note ===&lt;br /&gt;
IAN TO CHECK IN ACCESS&lt;br /&gt;
The reason the MS Access database seems to allow this condition is likely due to difference in character case between column names and primary key designations.&lt;br /&gt;
ERROR MESSAGE: Line 4: ERROR when executing SQL: column &amp;quot;OS_FOL_B_focal_AnimId&amp;quot; does not exist&lt;br /&gt;
Hint: Perhaps you meant to reference the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot; or the column &amp;quot;OTHER_SPECIES.OS_FOL_B_focal_AnimID&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== (#59) Zero BIOGRAPHY.b_animid_num values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
SokewDB requires that the animal ID number be greater than &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;, but some (what seem to be rows for babys) have a &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; value.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;select * from clean.biography where b_animid_num = 0;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Change the &amp;lt;code&amp;gt;0&amp;lt;/code&amp;gt; values to &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This is a brute-force, but adequate, solution because it does not validate anything concerning the rows affected.&lt;br /&gt;
&lt;br /&gt;
== * (#60) Invalid biography_update_log.made_by values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There is a &amp;lt;code&amp;gt;biography_update_log.made_by&amp;lt;/code&amp;gt; value (&amp;lt;code&amp;gt;SF/EVL&amp;lt;/code&amp;gt;) that is not a person.  (Not on the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; table.)&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.biography_update_log&lt;br /&gt;
  WHERE made_by IS NOT NULL&lt;br /&gt;
        AND NOT EXISTS (SELECT 1&lt;br /&gt;
                          FROM clean.people&lt;br /&gt;
                          WHERE people.person = biography_update_log.made_by);&amp;lt;/pre&amp;gt;&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN CHANGED THE SINGLE SF/EVL ENTRY TO EVL IN ACCESS 3/25/2026&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;update&amp;#039;&amp;#039;: this fix was not reflected in the 2026-07-12 dump.&lt;br /&gt;
&lt;br /&gt;
Suggest instead addressing this in clean:&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
-- Problem #60 BIOGRAPHY_UPDATE_LOG.MadeBy contains the invalid combination&lt;br /&gt;
-- person code SF/EVL.  The source-data resolution is to use EVL.&lt;br /&gt;
UPDATE biography_update_log&lt;br /&gt;
  SET made_by = &amp;#039;EVL&amp;#039;&lt;br /&gt;
  WHERE made_by = &amp;#039;SF/EVL&amp;#039;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#61) Invalid follow date/focals pairs in follow_arrival ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are &amp;lt;code&amp;gt;follow_arrival.fa_fol_b_animid&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;follow_arrival.fa_fol_date&amp;lt;/code&amp;gt; value combinations that do not exist in &amp;lt;code&amp;gt;follow&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
The invalid values can be listed (from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema) with:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM clean.follow_arrival&lt;br /&gt;
  WHERE NOT EXISTS&lt;br /&gt;
          (SELECT 1&lt;br /&gt;
             FROM clean.follow&lt;br /&gt;
             WHERE follow.fol_date = follow_arrival.fa_fol_date&lt;br /&gt;
                   AND follow.fol_b_animid&lt;br /&gt;
                   = follow_arrival.fa_fol_b_focal_animid)&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
===SOLUTION===&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== (#62) There are follow_arrival rows with arriving animids with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 39 follow arrivals where the arriving animid has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select * from easy.follow_arrival where fa_b_arr_animid &amp;lt;&amp;gt; rtrim(fa_b_arr_animid);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Remove the trailing spaces in table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 9c114ac.&lt;br /&gt;
&lt;br /&gt;
== * (#63) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_data_source values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 5,020 follow arrivals where the fa_data_source is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_data_source IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
Create a &amp;lt;code&amp;gt;none&amp;lt;/code&amp;gt; value in ARRIVAL_SOURCES, and use that value instead of &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; in follow_arrival table in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema. Addressed 8c05232.&lt;br /&gt;
&lt;br /&gt;
== (#64) There are follow_arrival rows with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; fa_type_of_certainty values ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is 1 follow arrivals where the fa_type_of_certainty is &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT *&lt;br /&gt;
  FROM easy.follow_arrival&lt;br /&gt;
  WHERE fa_type_of_certainty IS NULL&lt;br /&gt;
  ORDER BY follow_arrival.fa_fol_date&lt;br /&gt;
         , follow_arrival.fa_fol_b_focal_animid&lt;br /&gt;
         , follow_arrival.fa_b_arr_animid&lt;br /&gt;
         , follow_arrival.fa_time_start&lt;br /&gt;
         , follow_arrival.fa_time_end;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
IAN fixed in Access 3/25/2026&lt;br /&gt;
&lt;br /&gt;
== (#65) There are AGGRESSION_EVENT rows that have &amp;quot;YES&amp;quot; or a space as a ae_bad_observeration_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 3,079 rows where the ae_bad_observation_flag is a space, and 1 row where it is &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_bad_observation_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_bad_observation_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;YES&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;() as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Change the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema to contain only boolean values, based on the above.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#66) There are AGGRESSION_EVENT rows that have a space as a ae_decided_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,041 rows where the ae_decided_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_decided_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_decided_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
KARL: Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Stevan: Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#67) There are AGGRESSION_EVENT rows that have a space or an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; as a ae_multiple_aggressor_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,397 rows where the ae_multiple_aggressor_flag is a space and 2 rows where the value is an &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_aggressor_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_aggressor_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;x&amp;lt;/code&amp;gt; (along with &amp;lt;code&amp;gt;X&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;TRUE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== (#68) There are AGGRESSION_EVENT rows that have a space as a ae_multiple_recipient_flag value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 2,528 rows where the ae_multiple_recipient_flag is a space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;SELECT COALESCE(ae.ae_multiple_recipient_flag, &amp;#039;NULL&amp;#039;), count(*)&lt;br /&gt;
  FROM easy.aggression_event AS ae&lt;br /&gt;
  GROUP BY ae.ae_multiple_recipient_flag;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
Treat spaces (along with &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt;) as &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#69) There are duplicate year*community records in AGGRESSION_EVENT_LOG ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 11 duplicate pairs of year*community records in AGGRESSION_EVENT_LOG. There can be, at most, one row per-community, per-year. The duplicates each have a `b_rec_english` value of `ALL` or `F-F (ALL); F-M (ALL); M-M (ALL)`.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event_log&lt;br /&gt;
JOIN (&lt;br /&gt;
  SELECT&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  FROM&lt;br /&gt;
    clean.aggression_event_log&lt;br /&gt;
  GROUP BY&lt;br /&gt;
    year,&lt;br /&gt;
    community&lt;br /&gt;
  HAVING count(*) &amp;gt; 1&lt;br /&gt;
  ) AS dups ON (&lt;br /&gt;
  dups.year          = aggression_event_log.year&lt;br /&gt;
  AND dups.community = aggression_event_log.community&lt;br /&gt;
)&lt;br /&gt;
order by &lt;br /&gt;
  aggression_event_log.year,&lt;br /&gt;
  aggression_event_log.community ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#70) There are AGGRESSION_EVENT rows that do not have a matching FOLLOW for the same focal ID and date. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 1494 records where clean.aggression_event does not have a matching row in clean.follow for the same focal ID and date.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH problem_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ae.*,&lt;br /&gt;
        EXISTS (&lt;br /&gt;
            SELECT 1&lt;br /&gt;
            FROM clean.follow f_id&lt;br /&gt;
            WHERE f_id.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
        ) AS focal_id_exists_in_follow&lt;br /&gt;
    FROM clean.aggression_event ae&lt;br /&gt;
    WHERE NOT EXISTS (&lt;br /&gt;
        SELECT 1&lt;br /&gt;
        FROM clean.follow f&lt;br /&gt;
        WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
          AND f.fol_date = ae.ae_date&lt;br /&gt;
    )&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_time,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_b_aggressor_id,&lt;br /&gt;
    pr.ae_b_recipient_id,&lt;br /&gt;
    pr.ae_source,&lt;br /&gt;
    pr.ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN pr.focal_id_exists_in_follow THEN &amp;#039;missing_follow_on_same_date&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;focal_id_not_found_in_follow&amp;#039;&lt;br /&gt;
    END AS issue_type&lt;br /&gt;
FROM problem_rows pr&lt;br /&gt;
ORDER BY&lt;br /&gt;
    pr.ae_date,&lt;br /&gt;
    pr.ae_fol_b_focal_id,&lt;br /&gt;
    pr.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE SOLVED BY WATCHES TABLE&lt;br /&gt;
This means, aggressions can have a row on WATCHES, and go into the database, without there being a follow.&lt;br /&gt;
&lt;br /&gt;
I will mark resolved when I get the changes to aggression events done.  (Karl)&lt;br /&gt;
&lt;br /&gt;
== * (#71) There are AGGRESSION_EVENT animal ids that are not reflected in among FOLLOW animal ids. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 16 unique(!) aggression_event.ae_fol_b_focal_id records that are not reflected among follow.fol_b_animid values.&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
All of the problematic records are reflected in the rows excluded as part of Problem #70. As such, there is not a separate exclusion clause in the conversion process for these data. As a correlary, addressing all issues in Problem #70 would address all Problem #71 infractions well.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fol_b_focal_id NOT IN (&lt;br /&gt;
    SELECT DISTINCT fol_b_animid&lt;br /&gt;
    FROM clean.follow&lt;br /&gt;
    WHERE fol_b_animid IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
GROUP BY ae_fol_b_focal_id&lt;br /&gt;
ORDER BY&lt;br /&gt;
  row_count DESC,&lt;br /&gt;
  ae_fol_b_focal_id&lt;br /&gt;
;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN&lt;br /&gt;
&lt;br /&gt;
== (#72) There are AGGRESSION_EVENT recipient certainty flags other than `Y` or `N` (required). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 28,322 AGGRESSION_EVENT rows that have a ae_recipient_certainty_flag other than `Y` or `N` as required.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
ae.*,&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;) ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_flag,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM clean.follow f&lt;br /&gt;
  WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
  AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND UPPER(BTRIM(COALESCE(ae.ae_recipient_certainty_flag, &amp;#039;&amp;#039;))) NOT IN (&amp;#039;N&amp;#039;, &amp;#039;Y&amp;#039;)&lt;br /&gt;
GROUP BY COALESCE(ae_recipient_certainty_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;)&lt;br /&gt;
ORDER BY row_count DESC, raw_flag;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;N&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;NO&amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
** &amp;#039;Y%&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/2026&lt;br /&gt;
&lt;br /&gt;
KARL&lt;br /&gt;
&lt;br /&gt;
Stevan: The usual.  Fix in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
addressed as part of the solution to problem #74 (commit efa75af037a5e22b2b43675b435f472193a6672f)&lt;br /&gt;
&lt;br /&gt;
== * (#73) There are AGGRESSION_EVENT aggression_event.ae_b_recipient_id values that are NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
There are 5,820 AGGRESSION_EVENT rows for which the aggression_event.ae_b_recipient_id is NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
    ae.ae_recipient_certainty_flag,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NULL&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
CAN THESE ALL BE NULL? IF NOT, THEN WE NEED TO DO SOMETHING ABOUT INSTANCES WHEN &amp;#039;GROUP&amp;#039; IS THE TARGET.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
The usual solution is to have the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, be the target.  If NULL has some other special meaning, other than we don&amp;#039;t know at whom the aggression was directed, we could make up another &amp;quot;special&amp;quot; &amp;lt;code&amp;gt;BIOGRAPHY_DATA&amp;lt;/code&amp;gt; row, similar to the &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGF&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;MGM&amp;lt;/code&amp;gt;, etc. rows.&lt;br /&gt;
&lt;br /&gt;
Yes, use UNK.&lt;br /&gt;
&lt;br /&gt;
== (#74) There are AGGRESSION_EVENT event flags (multiple columns) that have values other than X or NULL. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
Among AGGRESSION_EVENT flags (ae_bad_observation_flag, ae_bristle_flag, ae_chase_flag, ae_contact_flag, ae_contact_flag, ae_decided_flag, ae_display_flag, ae_multiple_aggressor_flag, ae_multiple_recipient_flag, ae_vocal_flag, ae_vocal_flag) there are 2,864 rows that have a value other than `X` or NULL (required) for one or more of the flags.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  ae.ae_date,&lt;br /&gt;
  ae.ae_time,&lt;br /&gt;
  ae.ae_fol_b_focal_id,&lt;br /&gt;
  ae.ae_b_aggressor_id,&lt;br /&gt;
  ae.ae_b_recipient_id,&lt;br /&gt;
  ae.ae_decided_flag,&lt;br /&gt;
  ae.ae_multiple_aggressor_flag,&lt;br /&gt;
  ae.ae_multiple_recipient_flag,&lt;br /&gt;
  ae.ae_bad_observation_flag,&lt;br /&gt;
  ae.ae_bristle_flag,&lt;br /&gt;
  ae.ae_display_flag,&lt;br /&gt;
  ae.ae_chase_flag,&lt;br /&gt;
  ae.ae_contact_flag,&lt;br /&gt;
  ae.ae_vocal_flag,&lt;br /&gt;
  ae.ae_source,&lt;br /&gt;
  ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
)&lt;br /&gt;
AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
AND (&lt;br /&gt;
    (ae.ae_decided_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_decided_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_aggressor_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_aggressor_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_multiple_recipient_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_multiple_recipient_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bad_observation_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bad_observation_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_bristle_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_bristle_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_display_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_display_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_chase_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_chase_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_contact_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_contact_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
 OR (ae.ae_vocal_flag IS NOT NULL AND UPPER(BTRIM(ae.ae_vocal_flag)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;)&lt;br /&gt;
)&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH base AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
),&lt;br /&gt;
flag_values AS (&lt;br /&gt;
  SELECT &amp;#039;ae_decided_flag&amp;#039; AS flag_name, COALESCE(ae_decided_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_value FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_aggressor_flag&amp;#039;, COALESCE(ae_multiple_aggressor_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_multiple_recipient_flag&amp;#039;, COALESCE(ae_multiple_recipient_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bad_observation_flag&amp;#039;, COALESCE(ae_bad_observation_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_bristle_flag&amp;#039;, COALESCE(ae_bristle_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_display_flag&amp;#039;, COALESCE(ae_display_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_chase_flag&amp;#039;, COALESCE(ae_chase_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_contact_flag&amp;#039;, COALESCE(ae_contact_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT &amp;#039;ae_vocal_flag&amp;#039;, COALESCE(ae_vocal_flag, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) FROM base&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  flag_name,&lt;br /&gt;
  raw_value,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM flag_values&lt;br /&gt;
WHERE raw_value &amp;lt;&amp;gt; &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;&lt;br /&gt;
  AND UPPER(BTRIM(raw_value)) &amp;lt;&amp;gt; &amp;#039;X&amp;#039;&lt;br /&gt;
GROUP BY flag_name, raw_value&lt;br /&gt;
ORDER BY flag_name, row_count DESC, raw_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Ian or Elizabeth to please confirm.&lt;br /&gt;
&lt;br /&gt;
* clean schema conversions (mimic patterns in Access):&lt;br /&gt;
** `?` to &amp;#039; &amp;#039;&lt;br /&gt;
** `Y%` to `X`&lt;br /&gt;
** `X%` to `X`&lt;br /&gt;
&lt;br /&gt;
* sokwe conversions:&lt;br /&gt;
** &amp;#039; &amp;#039; to &amp;#039;0&amp;#039;&lt;br /&gt;
** &amp;#039;X&amp;#039; to &amp;#039;1&amp;#039;&lt;br /&gt;
&lt;br /&gt;
ICG AND ELO CONFIRM 7/3/26&lt;br /&gt;
&lt;br /&gt;
addressed with commit efa75af037a5e22b2b43675b435f472193a6672f&lt;br /&gt;
note that the solution to problem #74 was an umbrella fix that also addressed #65-68, and #72&lt;br /&gt;
&lt;br /&gt;
== * (#75) There are AGGRESSION_EVENT records where ae_extracted_by values are not in the people table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 28,513 AGGRESSION_EVENT records where ae_extracted_by does not match a person in the PEOPLE table.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  s.ae_date,&lt;br /&gt;
  s.ae_time,&lt;br /&gt;
  s.ae_fol_b_focal_id,&lt;br /&gt;
  s.ae_b_aggressor_id,&lt;br /&gt;
  s.ae_b_recipient_id,&lt;br /&gt;
  s.ae_extracted_by,&lt;br /&gt;
  s.ae_source,&lt;br /&gt;
  s.ae_full_description,&lt;br /&gt;
  s.ae_comments&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
ORDER BY s.ae_date, s.ae_fol_b_focal_id, s.ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH scoped AS (&lt;br /&gt;
  SELECT ae.*&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.follow f&lt;br /&gt;
    WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
      AND f.fol_date = ae.ae_date&lt;br /&gt;
  )&lt;br /&gt;
  AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;)) AS raw_extractedby,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM scoped s&lt;br /&gt;
WHERE NOT EXISTS (&lt;br /&gt;
  SELECT 1&lt;br /&gt;
  FROM people p&lt;br /&gt;
  WHERE p.person = BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
     OR LOWER(p.name) = LOWER(BTRIM(COALESCE(s.ae_extracted_by, &amp;#039;&amp;#039;)))&lt;br /&gt;
)&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_extracted_by, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_extractedby;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NOT SURE IF/HOW THE PEOPLE TABLE IS POPULATED.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
We populate the &amp;lt;code&amp;gt;PEOPLE&amp;lt;/code&amp;gt; however we like.  We just make things up.&lt;br /&gt;
We can add all these values to the table as different people.&lt;br /&gt;
However, unless you set the &amp;lt;code&amp;gt;PEOPLE.Active&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;FALSE&amp;lt;/code&amp;gt;&lt;br /&gt;
you run the risk of future data entry using a bunch of odd or outdated people.&lt;br /&gt;
&lt;br /&gt;
== (#76) There are AGGRESSION_EVENT rows where the Actor/Actee are not in the biography table. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,866 combined (but see following!) rows where the ae_b_aggressor_id and/or ae_b_recipient_id is not in the BIOGRAPHY table.&lt;br /&gt;
&lt;br /&gt;
Notes:&lt;br /&gt;
* The row count is inflated by a CROSS LATERAL JOIN, which yields the number of combined, pivoted records for both ae_b_aggressor_id and ae_b_recipient, not the actual number of confounding AGGRESSION_EVENT rows.&lt;br /&gt;
* The queries currently exclude ae_b_recipient_id values that are NULL, which seems a related but separte problem; the number of rows jumps to 10,357 if ae_b_recipient_id = NULL are included.&lt;br /&gt;
* How are we treating ae_b_recipient_id values such as `group`, `males`, `females`, etc.?&lt;br /&gt;
* Whitespace is considered in the incongruence such that, for example, `VIN ` (with a space) does not match `VIN` in biography.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae.ae_date,&lt;br /&gt;
    ae.ae_time,&lt;br /&gt;
    ae.ae_fol_b_focal_id,&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant,&lt;br /&gt;
    ae.ae_b_aggressor_id,&lt;br /&gt;
    ae.ae_b_recipient_id,&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
    ae.ae_source,&lt;br /&gt;
    ae.ae_full_description,&lt;br /&gt;
    ae.ae_comments&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae.ae_date, ae.ae_fol_b_focal_id, ae.ae_time, v.role_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;trimmed_match&amp;#039;&amp;#039; denotes a match between Actor/Actee and biography if whitespace is trimmed.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    v.role_name,&lt;br /&gt;
    v.participant AS raw_participant,&lt;br /&gt;
    b_trimmed.b_animid AS trimmed_match,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event ae&lt;br /&gt;
CROSS JOIN LATERAL (&lt;br /&gt;
  VALUES&lt;br /&gt;
    (&amp;#039;Actor&amp;#039;::text, ae.ae_b_aggressor_id),&lt;br /&gt;
    (&amp;#039;Actee&amp;#039;::text, ae.ae_b_recipient_id)&lt;br /&gt;
) AS v(role_name, participant)&lt;br /&gt;
LEFT JOIN clean.biography b_trimmed&lt;br /&gt;
  ON BTRIM(v.participant) = BTRIM(b_trimmed.b_animid)&lt;br /&gt;
WHERE v.participant IS NOT NULL&lt;br /&gt;
  AND v.participant &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM clean.biography b&lt;br /&gt;
    WHERE b.b_animid = v.participant&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY v.role_name, v.participant, b_trimmed.b_animid&lt;br /&gt;
ORDER BY row_count DESC, v.role_name, v.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ICG fixed in Access 7/7/2026. replaced &amp;quot;group&amp;quot;, &amp;quot;males&amp;quot;, etc with UNK. Fixed bad IDs. Updated comments field to reflect original entry terms.&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
&lt;br /&gt;
You can&amp;#039;t just &amp;quot;allow them&amp;quot;.  You can make rows in BIOGRAPHY_DATA that have AnimID values that correspond to the values that appear.  It&amp;#039;s hard to say whether or not these rows would have to be &amp;quot;special&amp;quot;.  Probably not, because anybody can aggress against anybody.  But if you&amp;#039;re going to make a whole bunch of &amp;quot;dummy&amp;quot; individuals (138 of them, to be exact), it might be prudent to add a flag to BIOGRAPHY_DATA that says whether or not the row corresponds to a specific chimpanzee.&lt;br /&gt;
(That flag could even be used to filter the dummy individuals out of the BIOGRAPHY view.)&lt;br /&gt;
&lt;br /&gt;
== * (#77) There are AGGRESSION_EVENT rows where ae_recipient_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 9,102 records where ae_recipient_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_recipient_behavior IS NULL&lt;br /&gt;
  OR ae_recipient_behavior = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#78) There are AGGRESSION_EVENT rows where ae_full_description is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 4,718 records where ae_full_description is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE&lt;br /&gt;
  ae_full_description IS NULL&lt;br /&gt;
  OR ae_full_description = &amp;#039;&amp;#039; ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Change the NULL values to the empty string in the clean schema.&lt;br /&gt;
&lt;br /&gt;
== (#79) There are AGGRESSION_EVENT participants outside their valid study participation window in the biography data. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 46 records where an Actor and/or Actee are in the AGGRESSION_EVENT table but outside their valid study participation window in the biography data. Note that the query strips white space when comparing the ids of aggression participant ids and biography animid ids given that the focus here is on checking for events outside of defined date ranges rather than matching ids and assuming that issues concerning white space will be resolved (also in the migration work-around). The query further winnows aggression events for which there is a valid follow record (not in the migration work-around).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH candidate_agg AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;)) AS actor_id,&lt;br /&gt;
      BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;)) AS actee_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.follow f&lt;br /&gt;
          WHERE f.fol_b_animid = ae.ae_fol_b_focal_id&lt;br /&gt;
            AND f.fol_date     = ae.ae_date&lt;br /&gt;
        )&lt;br /&gt;
    AND ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_aggressor_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
    AND EXISTS (&lt;br /&gt;
          SELECT 1&lt;br /&gt;
          FROM clean.biography b&lt;br /&gt;
          WHERE b.b_animid = BTRIM(COALESCE(ae.ae_b_recipient_id, &amp;#039;&amp;#039;))&lt;br /&gt;
        )&lt;br /&gt;
),&lt;br /&gt;
participants AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actor&amp;#039; AS role,&lt;br /&gt;
      c.actor_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
  UNION ALL&lt;br /&gt;
  SELECT&lt;br /&gt;
      c.ae_date,&lt;br /&gt;
      c.ae_time,&lt;br /&gt;
      c.ae_fol_b_focal_id,&lt;br /&gt;
      &amp;#039;Actee&amp;#039; AS role,&lt;br /&gt;
      c.actee_id AS participant,&lt;br /&gt;
      c.ae_source,&lt;br /&gt;
      c.ae_full_description&lt;br /&gt;
  FROM candidate_agg c&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    p.ae_date,&lt;br /&gt;
    p.ae_time,&lt;br /&gt;
    p.ae_fol_b_focal_id,&lt;br /&gt;
    p.role,&lt;br /&gt;
    p.participant,&lt;br /&gt;
    b.b_entrydate,&lt;br /&gt;
    b.b_departdate,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN p.ae_date &amp;lt; b.b_entrydate THEN &amp;#039;before_entry&amp;#039;&lt;br /&gt;
      WHEN p.ae_date &amp;gt; b.b_departdate THEN &amp;#039;after_departure&amp;#039;&lt;br /&gt;
      ELSE &amp;#039;ok&amp;#039;&lt;br /&gt;
    END AS violation_type,&lt;br /&gt;
    p.ae_source,&lt;br /&gt;
    p.ae_full_description&lt;br /&gt;
FROM participants p&lt;br /&gt;
JOIN clean.biography b&lt;br /&gt;
  ON b.b_animid = p.participant&lt;br /&gt;
WHERE p.ae_date &amp;lt; b.b_entrydate&lt;br /&gt;
   OR p.ae_date &amp;gt; b.b_departdate&lt;br /&gt;
ORDER BY p.ae_date, p.ae_fol_b_focal_id, p.ae_time, p.role, p.participant;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ICG FIXED IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
==== Karl comment ====&lt;br /&gt;
There is no good way to &amp;quot;allow for now&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
We could change the under study dates (EntryDate, DepartDate) of the affected individuals &amp;quot;for now&amp;quot;.  That would &amp;quot;solve&amp;quot; the problem.&lt;br /&gt;
&lt;br /&gt;
Or just don&amp;#039;t convert the data and let the users get the info from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema and integrate it into their processing.&lt;br /&gt;
&lt;br /&gt;
== * (#80) There are AGGRESSION_EVENT rows where ae_aggressor_behavior is NULL or an empty string. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 143 records where ae_aggressor_behavior is NULL or an empty string.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_aggressor_behavior IS NULL&lt;br /&gt;
   OR BTRIM(ae_aggressor_behavior) = &amp;#039;&amp;#039;&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN POPULATE USING THE FULL DESCRIPTION FIELD.&lt;br /&gt;
&lt;br /&gt;
Change the &amp;lt;code&amp;gt;NULL&amp;lt;/code&amp;gt; values to the empty string in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#81) There are AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;). ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 112 AGGRESSION_EVENT rows where ae_time is NULL or outside the allowable window (&amp;#039;04:00:00&amp;#039;, &amp;#039;20:00:00&amp;#039;).&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    CASE&lt;br /&gt;
      WHEN ae_time IS NULL THEN &amp;#039;null_time&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;below_min&amp;#039;&lt;br /&gt;
      WHEN ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;above_max&amp;#039;&lt;br /&gt;
    END AS time_issue&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_time IS NULL&lt;br /&gt;
   OR ae_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR ae_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
MOST ARE NULL VALUES BECAUSE THEY WERE &amp;#039;BOOK EXTRACTS&amp;#039;. ASSIGN A &amp;#039;DUMMY&amp;#039; TIME?&lt;br /&gt;
&lt;br /&gt;
ICG FIXED THE FEW THAT WERE NOT BETWEEN 4:00 AND 20:00 IN ACCESS 7/6/2026&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Allowing means changing the &amp;quot;too early&amp;quot; and &amp;quot;too late&amp;quot; limits for all times, appearing anywhere in the db.  This is easy, but opens the door for other data errors.&lt;br /&gt;
&lt;br /&gt;
The change to the time limits cannnot (easily) be restricted to just the table that is converted.  Well.... It can.  But the change will not be reflected in the documentation.&lt;br /&gt;
(And it&amp;#039;ll be completely annoying because a change will have to be re-made each time the conversion program is re-run, which is a lot because we have a lot of errors to resolve.)&lt;br /&gt;
&lt;br /&gt;
Changing the limits back is less than easy.  A separate SQL statement needs to be run for each time column that appears in the database.  Probably better is to dump the database content, rebuild the database tables, etc., and then re-load the database content.&lt;br /&gt;
If any time values are out of bounds, the re-load of the database content will fail.&lt;br /&gt;
&lt;br /&gt;
Executive Summary: Anything&amp;#039;s possible but you&amp;#039;re coloring outside the lines here.&lt;br /&gt;
&lt;br /&gt;
== * (#82) There are AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,041 AGGRESSION_EVENT rows where severity (ae_fight_category) is not an allowable value [(&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)]. Note that all non-compliant values are NULL or some variation of white space.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_fight_category,&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS normalized_fight_category,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;)) AS offending_value,&lt;br /&gt;
    COUNT(*) AS row_count&lt;br /&gt;
FROM clean.aggression_event&lt;br /&gt;
WHERE ae_fight_category IS NOT NULL&lt;br /&gt;
  AND (&lt;br /&gt;
       BTRIM(ae_fight_category) = &amp;#039;&amp;#039;&lt;br /&gt;
       OR BTRIM(ae_fight_category) NOT IN (&amp;#039;unrated&amp;#039;, &amp;#039;0&amp;#039;, &amp;#039;1&amp;#039;, &amp;#039;2&amp;#039;, &amp;#039;?&amp;#039;)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(COALESCE(ae_fight_category, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, offending_value;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Fine-grain assessment of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Note the &amp;#039;&amp;#039;white space&amp;#039;&amp;#039; problem.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select distinct &amp;#039;&amp;quot;&amp;#039; || ae_fight_category || &amp;#039;&amp;quot;&amp;#039; from clean.aggression_event ;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW - MOST DATA ENTERERS DIDN&amp;#039;T FILL THIS COLUMN IN&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
There seems to be only a problem because NULL values exist.  The other non-conformant data seems to be gone from MS Access.  So make a fight category code that is &amp;quot;no data&amp;quot; and clean up the data in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
NULL means unrated&lt;br /&gt;
&lt;br /&gt;
== * (#83) There are AGGRESSION_EVENT rows where the Actor and Actee have the same animal ID. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 1,832 &amp;lt;s&amp;gt;363&amp;lt;/s&amp;gt; AGGRESSION_EVENT rows where `ae_b_aggressor_id` and `ae_b_recipient_id` share the same animal id.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH potential_role_dupes AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
      ae.ae_date,&lt;br /&gt;
      ae.ae_time,&lt;br /&gt;
      ae.ae_fol_b_focal_id,&lt;br /&gt;
      ae.ae_b_aggressor_id,&lt;br /&gt;
      ae.ae_b_recipient_id,&lt;br /&gt;
      ae.ae_source,&lt;br /&gt;
      ae.ae_full_description,&lt;br /&gt;
      ae.ae_comments,&lt;br /&gt;
      ae.dup&lt;br /&gt;
  FROM clean.aggression_event ae&lt;br /&gt;
  WHERE ae.ae_b_recipient_id IS NOT NULL&lt;br /&gt;
    AND ae.ae_b_aggressor_id IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    ae_date,&lt;br /&gt;
    ae_time,&lt;br /&gt;
    ae_fol_b_focal_id,&lt;br /&gt;
    ae_b_aggressor_id,&lt;br /&gt;
    ae_b_recipient_id,&lt;br /&gt;
    ae_source,&lt;br /&gt;
    ae_full_description,&lt;br /&gt;
    ae_comments,&lt;br /&gt;
    dup&lt;br /&gt;
FROM potential_role_dupes&lt;br /&gt;
WHERE ae_b_aggressor_id = ae_b_recipient_id&lt;br /&gt;
ORDER BY ae_date, ae_fol_b_focal_id, ae_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Processing note ===&lt;br /&gt;
&lt;br /&gt;
This data problem generates the following error but note that, in practice, this error could be generated for other reasons as well.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
./load_chunks.sh load_aggressions.m4 clean.aggression_event&lt;br /&gt;
psql:&amp;lt;stdin&amp;gt;:394: ERROR:  duplicate key value violates unique constraint &amp;quot;On ROLES, Participant + EID must be unique&amp;quot;&lt;br /&gt;
DETAIL:  Key (participant, eid)=(FD, 759513) already exists.&lt;br /&gt;
CONTEXT:  SQL statement &amp;quot;INSERT INTO roles (&lt;br /&gt;
        eid&lt;br /&gt;
      , role&lt;br /&gt;
      , participant)&lt;br /&gt;
    VALUES (&lt;br /&gt;
      CURRVAL(&amp;#039;events_eid_seq&amp;#039;)&lt;br /&gt;
    , &amp;#039;Actee&amp;#039;&lt;br /&gt;
    , this_ae.ae_b_recipient_id)&amp;quot;&lt;br /&gt;
PL/pgSQL function inline_code_block line 186 at SQL statement&lt;br /&gt;
make: *** [Makefile:373: load_aggressions] Error 3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. EVENTUALLY A VOLUNTEER CAN CHECK AND FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
If we allow, we allow for all event types, not just aggressions.  Groomings, matings, etc would all allow &amp;quot;self-dealing&amp;quot;.  I&amp;#039;d rather not.&lt;br /&gt;
&lt;br /&gt;
How about not converting this data and letting people who want it get it from the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
do not convert these rows, fix them later&lt;br /&gt;
&lt;br /&gt;
== * (#84) Some follows have a community with trailing spaces ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 3 follows where the fol_cl_community_id has trailing spaces.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
Cleaned in &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema, so query &amp;lt;code&amp;gt;easy&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
select &amp;#039;&amp;quot;&amp;#039; || fol_cl_community_id || &amp;#039;&amp;quot;&amp;#039; AS fol_cl_community_id_untrimmed,&lt;br /&gt;
*&lt;br /&gt;
FROM easy.follow&lt;br /&gt;
WHERE RTRIM(fol_cl_community_id) &amp;lt;&amp;gt; fol_cl_community_id;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
KARL TO FIX&lt;br /&gt;
&lt;br /&gt;
Fix by trimming spaces in the &amp;lt;code&amp;gt;clean&amp;lt;/code&amp;gt; schema.&lt;br /&gt;
&lt;br /&gt;
== * (#85) There are numerous instances of local food names that translate to multiple scientific food names. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 47 records documenting instances where a local food name (fl_local_food_name) is associated with more than one food scientific food name (fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH duplicates AS (&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    COUNT(*) AS count&lt;br /&gt;
  FROM clean.food_lookup&lt;br /&gt;
  GROUP BY fl_sci_food_name&lt;br /&gt;
  HAVING COUNT(*) &amp;gt; 1&lt;br /&gt;
  )&lt;br /&gt;
  SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    duplicates.fl_sci_food_name&lt;br /&gt;
FROM clean.food_lookup&lt;br /&gt;
JOIN duplicates ON clean.food_lookup.fl_sci_food_name = duplicates.fl_sci_food_name&lt;br /&gt;
ORDER BY fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #87. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name (i.e., not fl_sci_food_name_gen); else, refer to #87.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW, BUT NEEDS TO BE DISCUSSED AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
To allow, do some hackery when constructing the support tables.  Make non-unique values unique by adding extra &amp;quot;duplicate #1 --&amp;quot; sort of text into the value of the column.  Either &amp;quot;by hand&amp;quot; in the conversion code or by way of some algorythm.&lt;br /&gt;
&lt;br /&gt;
== * (#86) There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are numerous instances in the food_part_lookup that that conflate name, initial, and/or the english translation. These issues are detailed below. Please note that this problem concerns the food_part_lookup table specifically.&lt;br /&gt;
&lt;br /&gt;
Related, we need to clarify how we are treating the initials of food parts. Currently, initials are added to the codes.food_parts table with the corresponding food part name but with a trailing `-initial` so as not to violate the &lt;br /&gt;
codes.food_parts.description unique constraint. So, both the food part name and food part initials are included in food events. This is probably not a good approach. Better would be to map the initial to the corresponding food part name and use only the full name (but then you do lose some reference to the original data). Ian needs to please consider.&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
==== query ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH source_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        ROW_NUMBER() OVER () AS src_ord,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_local_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS local_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_food_part_initials, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS initials_arr,&lt;br /&gt;
&lt;br /&gt;
        COALESCE((&lt;br /&gt;
            SELECT ARRAY_AGG(tok)&lt;br /&gt;
            FROM (&lt;br /&gt;
                SELECT BTRIM(x) AS tok&lt;br /&gt;
                FROM REGEXP_SPLIT_TO_TABLE(COALESCE(fpl_english_food_part, &amp;#039;&amp;#039;), E&amp;#039;[;:,]&amp;#039;) AS x&lt;br /&gt;
                WHERE BTRIM(x) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
            ) s&lt;br /&gt;
        ), ARRAY[]::text[]) AS english_arr&lt;br /&gt;
&lt;br /&gt;
    FROM clean.food_part_lookup&lt;br /&gt;
),&lt;br /&gt;
expanded AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        s.src_ord,&lt;br /&gt;
        gs.idx,&lt;br /&gt;
        COALESCE(s.local_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_local_food_part,&lt;br /&gt;
        COALESCE(s.initials_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_food_part_initials,&lt;br /&gt;
        COALESCE(s.english_arr[gs.idx], &amp;#039;&amp;#039;) AS fpl_english_food_part&lt;br /&gt;
    FROM source_rows s&lt;br /&gt;
    CROSS JOIN LATERAL GENERATE_SERIES(&lt;br /&gt;
        1,&lt;br /&gt;
        GREATEST(&lt;br /&gt;
            CARDINALITY(s.local_arr),&lt;br /&gt;
            CARDINALITY(s.initials_arr),&lt;br /&gt;
            CARDINALITY(s.english_arr)&lt;br /&gt;
        )&lt;br /&gt;
    ) AS gs(idx)&lt;br /&gt;
),&lt;br /&gt;
deduped AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        e.*,&lt;br /&gt;
        ROW_NUMBER() OVER (&lt;br /&gt;
            PARTITION BY&lt;br /&gt;
                e.fpl_local_food_part,&lt;br /&gt;
                e.fpl_food_part_initials,&lt;br /&gt;
                e.fpl_english_food_part&lt;br /&gt;
            ORDER BY e.src_ord, e.idx&lt;br /&gt;
        ) AS rn&lt;br /&gt;
    FROM expanded e&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part&lt;br /&gt;
FROM deduped&lt;br /&gt;
WHERE rn = 1&lt;br /&gt;
ORDER BY src_ord, idx&lt;br /&gt;
    fpl_local_food_part,&lt;br /&gt;
    fpl_food_part_initials,&lt;br /&gt;
    fpl_english_food_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== summary ====&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! fpl_local_food_part !! fpl_food_part_initials !! fpl_english_food_part&lt;br /&gt;
|-&lt;br /&gt;
| CHIPUKIZA      || C  || SHOOTS&lt;br /&gt;
|-&lt;br /&gt;
| MAJANI         || J  || LEAVES&lt;br /&gt;
|-&lt;br /&gt;
| MBEGU          || MB || SEEDS&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || W  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| MABUA          || B  || PITH&lt;br /&gt;
|-&lt;br /&gt;
| MAGOMA         || G  || BARK&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA        || T  || FRUIT&lt;br /&gt;
|-&lt;br /&gt;
| MAUA           || M  || FLOWERS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVI         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU WENGINE || D  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UTOMVU         || U  || SAP&lt;br /&gt;
|-&lt;br /&gt;
| MCHWA          || W  || TERMITES&lt;br /&gt;
|-&lt;br /&gt;
| MIFUPA         || NA || BONES&lt;br /&gt;
|-&lt;br /&gt;
| MITI           || NA || TREE&lt;br /&gt;
|-&lt;br /&gt;
| MIZIZI         || NA || ROOTS&lt;br /&gt;
|-&lt;br /&gt;
| NA             || NA || NOT APPLICABLE&lt;br /&gt;
|-&lt;br /&gt;
| NONE           || NA || None&lt;br /&gt;
|-&lt;br /&gt;
| NYAMA          || N  || MEAT&lt;br /&gt;
|-&lt;br /&gt;
| SIAFU          || S  || INSECTS&lt;br /&gt;
|-&lt;br /&gt;
| UNRECORDED     || NA || UNRECORDED&lt;br /&gt;
|-&lt;br /&gt;
| WADUDU         || D  || INSECTS&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== notes ====&lt;br /&gt;
&lt;br /&gt;
* there are two initials (`W`, `D`) for `WADUDU` ~ `INSECTS`&lt;br /&gt;
* need to clarify `WADUDU WENGINE`, which also shares an initial (`D`) and english translation (`INSECTS`) as `WADUDU`&lt;br /&gt;
* the initial `W` is associated with both `WADUDU` and `MCHWA`&lt;br /&gt;
* different spellings for `SAP`: `UTOMVU` and `UTOMVI`&lt;br /&gt;
* `INSECTS` associated with `WADUDU`, `WADUDU WENGINE`, and `SIAFU`&lt;br /&gt;
* how should we treat `NA`, `NONE`, and `UNRECORDED`&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. IAN/PI GROUP TO DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
== * (#87) There are numerous duplicate fl_sci_food_name_gen values in food_lookup. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 42 instances where fl_sci_food_name_gen values are associated with more than one fl_local_food_name value. This problem assumes that we are using the fl_sci_food_name_gen value for food_lookup.description (i.e., instead of fl_sci_food_name). If instead, food_lookup.description should reflect fl_sci_food_name then this particular problem is moot and can be ignored (but other problems will certainly arise with the switch to fl_sci_food_name).&lt;br /&gt;
&lt;br /&gt;
=== Bad Data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH ranked AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fl_local_food_name,&lt;br /&gt;
        fl_sci_food_name,&lt;br /&gt;
        fl_sci_food_name_gen,&lt;br /&gt;
        COUNT(*) OVER (&lt;br /&gt;
            PARTITION BY BTRIM(fl_sci_food_name_gen)&lt;br /&gt;
        ) AS sci_gen_count&lt;br /&gt;
    FROM clean.food_lookup&lt;br /&gt;
    WHERE fl_sci_food_name_gen IS NOT NULL&lt;br /&gt;
      AND BTRIM(fl_sci_food_name_gen) &amp;lt;&amp;gt; &amp;#039;&amp;#039;&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fl_local_food_name,&lt;br /&gt;
    fl_sci_food_name,&lt;br /&gt;
    fl_sci_food_name_gen,&lt;br /&gt;
    sci_gen_count&lt;br /&gt;
FROM ranked&lt;br /&gt;
WHERE sci_gen_count &amp;gt; 1&lt;br /&gt;
ORDER BY fl_sci_food_name_gen, fl_local_food_name, fl_sci_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Notes ===&lt;br /&gt;
&lt;br /&gt;
But see problem #85. This problem is only relevant if codes.food_names.description is derived from fl_sci_food_name_gen (i.e., not fl_sci_food_name); else, refer to #85.&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW, DISCUSS AS A PI GROUP&lt;br /&gt;
&lt;br /&gt;
== * (#88) There are food_bout rows for which there is not a corresponding food name. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 2,114 clean.food_bout rows for which there is not a matching codes.food_names.name.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
Exact-match diagnostic: food_bout food names missing from codes.food_names&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
Use this to list only the distinct food names present in clean.food_bout but missing from codes.food_names.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== white-space specific mismatches ====&lt;br /&gt;
&lt;br /&gt;
Whitespace-only diagnostic: rows excluded only because of whitespace differences This finds values whose trimmed form exists in codes.food_names, but the raw value does not.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH food_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        1 AS seq,&lt;br /&gt;
        fb_fl_local_food_name AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_fl_local_food_name) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE fb_fl_local_food_name IS NOT NULL&lt;br /&gt;
    UNION ALL&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb_fol_date,&lt;br /&gt;
        fb_fol_b_focal_animid,&lt;br /&gt;
        fb_begin_feed_time,&lt;br /&gt;
        fb_end_feed_time,&lt;br /&gt;
        2 AS seq,&lt;br /&gt;
        fb_local_food_name2 AS raw_foodname,&lt;br /&gt;
        BTRIM(fb_local_food_name2) AS trimmed_foodname&lt;br /&gt;
    FROM clean.food_bout&lt;br /&gt;
    -- WHERE NULLIF(fb_local_food_name2, &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname,&lt;br /&gt;
    trimmed_foodname&lt;br /&gt;
FROM food_rows&lt;br /&gt;
WHERE raw_foodname NOT IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
AND trimmed_foodname IN (&lt;br /&gt;
    SELECT name&lt;br /&gt;
    FROM codes.food_names&lt;br /&gt;
)&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    seq,&lt;br /&gt;
    raw_foodname;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO CHECK WHETHER LEGIT NEW FOODS OR JUST POOR HANDWRITING, SPELLING, ETC&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Ok.  Allowing means adding a bunch of odd values, that could continue to be used in the future until the data is cleaned and the &amp;quot;extra&amp;quot; codes removed.&lt;br /&gt;
&lt;br /&gt;
== * (#89) There is a food_bout row that has both a compound fb_fpl_local_food_part value and where fb_fpl_local_food_part2 is not null. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! FB_FPL_local_food_part !! FB_FPL_local_food_part2&lt;br /&gt;
|-&lt;br /&gt;
| MATUNDA; CHIPUKIZA || MATUNDA; CHIPUKIZA&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
SHOULD BE ALLOWED&lt;br /&gt;
&lt;br /&gt;
== * (#90) There are food_bout rows for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 229 records for which a food_bout row does not have a match between for which the fb_fpl_local_food_part or fb_fpl_local_food_part2 and codes.food_parts. Note, however, that the record count is somewhat ambiguous since we are deailing with multiple parts (1 and 2), and multiple rows per event.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
diagnostic: fb_fpl_local_food_part or fb_fpl_local_food_part2 does not have a matching value in codes.food_parts&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    seq1_part,&lt;br /&gt;
    seq2_part,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq1_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
        WHEN seq2_part IS NOT NULL&lt;br /&gt;
             AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
        THEN &amp;#039;seq2_part missing from codes.food_parts&amp;#039;&lt;br /&gt;
    END AS exclusion_reason&lt;br /&gt;
FROM prepared&lt;br /&gt;
WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
   OR (&lt;br /&gt;
        seq2_part IS NOT NULL&lt;br /&gt;
        AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
      )&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        1&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq1_part,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN regexp_replace(&lt;br /&gt;
                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                     &amp;#039;g&amp;#039;&lt;br /&gt;
                 ) LIKE &amp;#039;%;%&amp;#039;&lt;br /&gt;
            THEN UPPER(&lt;br /&gt;
                     NULLIF(&lt;br /&gt;
                         BTRIM(&lt;br /&gt;
                             split_part(&lt;br /&gt;
                                 regexp_replace(&lt;br /&gt;
                                     COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                                     E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                                     &amp;#039;;&amp;#039;,&lt;br /&gt;
                                     &amp;#039;g&amp;#039;&lt;br /&gt;
                                 ),&lt;br /&gt;
                                 &amp;#039;;&amp;#039;,&lt;br /&gt;
                                 2&lt;br /&gt;
                             )&lt;br /&gt;
                         ),&lt;br /&gt;
                         &amp;#039;&amp;#039;&lt;br /&gt;
                     )&lt;br /&gt;
                 )&lt;br /&gt;
            ELSE UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;))&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
)&lt;br /&gt;
SELECT missing_part, COUNT(*) AS occurrences&lt;br /&gt;
FROM (&lt;br /&gt;
    SELECT seq1_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq1_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
&lt;br /&gt;
    UNION ALL&lt;br /&gt;
&lt;br /&gt;
    SELECT seq2_part AS missing_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
      AND seq2_part NOT IN (SELECT part FROM codes.food_parts)&lt;br /&gt;
) x&lt;br /&gt;
GROUP BY missing_part&lt;br /&gt;
ORDER BY missing_part;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
ALLOW FOR NOW. NEED TO HAND CHECK AND DECIDE ON PROTOCOL&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
Stevan can chime in here, but I think this means adding more illegitimate &amp;quot;legitimate&amp;quot; codes.&lt;br /&gt;
&lt;br /&gt;
== * (#91) There are food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 43 clean.food_bout rows that have fb_begin_feed_time &amp;gt; fb_end_feed_time.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
IAN/VOLUNTEER TO FIX&lt;br /&gt;
&lt;br /&gt;
== * (#92) There are food_bout rows that do not have a corresponding follow. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
clean.food_bout references follows by (fb_fol_b_focal_animid, fb_fol_date), but not all rows (n=455) match sokwedb.follows exactly.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
Includes both:&lt;br /&gt;
* Normalization-only matches: match after TRIM+UPPER (raw focal text quality issue).&lt;br /&gt;
* True misses: no matching follow even after normalization (actual source data gap).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    fb_begin_feed_time - fb_end_feed_time AS inverted_by,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fl_local_food_name&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_end_feed_time IS NOT NULL&lt;br /&gt;
  AND fb_begin_feed_time &amp;gt; fb_end_feed_time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
NEED TO LOOK AT THE SOURCE OF THE PROBLEM B/C NO-ONE ENTERS FEEDING FROM B-REC NOTES.&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comments ====&lt;br /&gt;
&lt;br /&gt;
I think it makes sense to normalize the data in the clean schema, upper casing and removing spaces.  And then deal with what&amp;#039;s left.&lt;br /&gt;
&lt;br /&gt;
== * (#93) Need to address how we treat behaviours for the second value of food parts AND how they are derived. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
1. In load_food_events.sql, we are loading into table food_events, where each row requires both:&lt;br /&gt;
&lt;br /&gt;
* foodpart (NOT NULL + FK)&lt;br /&gt;
* foodname (NOT NULL + FK)&lt;br /&gt;
&lt;br /&gt;
...but for seq=2, these cases behave differently:&lt;br /&gt;
&lt;br /&gt;
# part2 exists, name2 missing: insert NODATA to name2&lt;br /&gt;
# part2 missing, name2 exists: insert NODATA to part2&lt;br /&gt;
# both missing: no seq=2 row.&lt;br /&gt;
# both exist: normal seq=2 row.&lt;br /&gt;
&lt;br /&gt;
This approach required adding NODATA to codes.food_names and codes.food_parts.&lt;br /&gt;
&lt;br /&gt;
Require approval for this approach or an altnernative.&lt;br /&gt;
&lt;br /&gt;
2. Multiple food parts can be dervied in one of two ways:&lt;br /&gt;
&lt;br /&gt;
# there are values for both fb_fpl_local_food_part and fb_fpl_local_food_part2, or&lt;br /&gt;
# there are multiple food parts in fb_fpl_local_food_part separated by a `;` (usually) or, rarely, `:`&lt;br /&gt;
&lt;br /&gt;
Do we need to document or log if fb_fpl_local_food_part2 is derived from a compound fb_fpl_local_food_part value?&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== count of values that meet these conditions ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
),&lt;br /&gt;
seq2_rows AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
        COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
    FROM effective&lt;br /&gt;
    WHERE seq2_part IS NOT NULL&lt;br /&gt;
       OR seq2_name IS NOT NULL&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; AND final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;both_nodata&amp;#039;&lt;br /&gt;
        WHEN final_seq2_part = &amp;#039;NODATA&amp;#039; THEN &amp;#039;part_nodata_only&amp;#039;&lt;br /&gt;
        WHEN final_seq2_name = &amp;#039;NODATA&amp;#039; THEN &amp;#039;name_nodata_only&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;no_nodata&amp;#039;&lt;br /&gt;
    END AS seq2_case,&lt;br /&gt;
    COUNT(*) AS rows&lt;br /&gt;
FROM seq2_rows&lt;br /&gt;
GROUP BY 1&lt;br /&gt;
ORDER BY 1;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
WITH prepared AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        fb.*,&lt;br /&gt;
        regexp_replace(&lt;br /&gt;
            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
            &amp;#039;;&amp;#039;,&lt;br /&gt;
            &amp;#039;g&amp;#039;&lt;br /&gt;
        ) AS part1_norm,&lt;br /&gt;
        UPPER(&lt;br /&gt;
            NULLIF(&lt;br /&gt;
                BTRIM(&lt;br /&gt;
                    split_part(&lt;br /&gt;
                        regexp_replace(&lt;br /&gt;
                            COALESCE(fb.fb_fpl_local_food_part, &amp;#039;&amp;#039;),&lt;br /&gt;
                            E&amp;#039;\\s*[:;]\\s*&amp;#039;,&lt;br /&gt;
                            &amp;#039;;&amp;#039;,&lt;br /&gt;
                            &amp;#039;g&amp;#039;&lt;br /&gt;
                        ),&lt;br /&gt;
                        &amp;#039;;&amp;#039;,&lt;br /&gt;
                        2&lt;br /&gt;
                    )&lt;br /&gt;
                ),&lt;br /&gt;
                &amp;#039;&amp;#039;&lt;br /&gt;
            )&lt;br /&gt;
        ) AS seq2_from_part1,&lt;br /&gt;
        UPPER(NULLIF(BTRIM(fb.fb_fpl_local_food_part2), &amp;#039;&amp;#039;)) AS seq2_from_part2,&lt;br /&gt;
        NULLIF(BTRIM(fb.fb_local_food_name2), &amp;#039;&amp;#039;) AS seq2_name&lt;br /&gt;
    FROM clean.food_bout fb&lt;br /&gt;
),&lt;br /&gt;
effective AS (&lt;br /&gt;
    SELECT&lt;br /&gt;
        *,&lt;br /&gt;
        CASE&lt;br /&gt;
            WHEN part1_norm LIKE &amp;#039;%;%&amp;#039; THEN seq2_from_part1&lt;br /&gt;
            ELSE seq2_from_part2&lt;br /&gt;
        END AS seq2_part&lt;br /&gt;
    FROM prepared&lt;br /&gt;
)&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fpl_local_food_part,&lt;br /&gt;
    fb_fpl_local_food_part2,&lt;br /&gt;
    fb_fl_local_food_name,&lt;br /&gt;
    fb_local_food_name2,&lt;br /&gt;
    COALESCE(seq2_part, &amp;#039;NODATA&amp;#039;) AS final_seq2_part,&lt;br /&gt;
    COALESCE(seq2_name, &amp;#039;NODATA&amp;#039;) AS final_seq2_name&lt;br /&gt;
FROM effective&lt;br /&gt;
WHERE seq2_part IS NOT NULL&lt;br /&gt;
   OR seq2_name IS NOT NULL&lt;br /&gt;
ORDER BY&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_fl_local_food_name;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
GET STEVAN TO EXPLAIN TO ICG AND ELO&lt;br /&gt;
&lt;br /&gt;
== * (#94) There are food_bout rows where the start or end of the observation is outside of allowable bounds. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 11 record where the food_bout is is outside of allowable bounds.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_begin_feed_time IS NULL THEN &amp;#039;BEGIN_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;BEGIN_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;BEGIN_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;BEGIN_OK&amp;#039;&lt;br /&gt;
    END AS begin_status,&lt;br /&gt;
    CASE&lt;br /&gt;
        WHEN fb_end_feed_time IS NULL THEN &amp;#039;END_NULL&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time THEN &amp;#039;END_BEFORE_MIN&amp;#039;&lt;br /&gt;
        WHEN fb_end_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time THEN &amp;#039;END_AFTER_MAX&amp;#039;&lt;br /&gt;
        ELSE &amp;#039;END_OK&amp;#039;&lt;br /&gt;
    END AS end_status,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_begin_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_begin,&lt;br /&gt;
    COALESCE(&lt;br /&gt;
        LEAST(GREATEST(fb_end_feed_time, &amp;#039;04:00:00&amp;#039;::time), &amp;#039;20:00:00&amp;#039;::time),&lt;br /&gt;
        &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
    ) AS clamped_end&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE fb_begin_feed_time IS NULL&lt;br /&gt;
   OR fb_end_feed_time IS NULL&lt;br /&gt;
   OR fb_begin_feed_time &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_begin_feed_time &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;lt; &amp;#039;04:00:00&amp;#039;::time&lt;br /&gt;
   OR fb_end_feed_time   &amp;gt; &amp;#039;20:00:00&amp;#039;::time&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
ALLOW FOR NOW. IAN TO FIX&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
See remarks on problem #81.&lt;br /&gt;
&lt;br /&gt;
== (#95) There is at least one food_bout rows where the start or stop time has a seconds value. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is one food_bout row where the start or stop time has a seconds value, which violates our time-check constraints.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    fb_fol_date,&lt;br /&gt;
    fb_fol_b_focal_animid,&lt;br /&gt;
    fb_begin_feed_time,&lt;br /&gt;
    fb_end_feed_time&lt;br /&gt;
FROM clean.food_bout&lt;br /&gt;
WHERE EXTRACT(SECOND FROM fb_begin_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM fb_end_feed_time) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY fb_fol_date, fb_fol_b_focal_animid, fb_begin_feed_time;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== SOLUTION ===&lt;br /&gt;
&lt;br /&gt;
IAN (HOPEFULLY) FIXED IN ACCESS 7/3&lt;br /&gt;
&lt;br /&gt;
==== Karl&amp;#039;s comment ====&lt;br /&gt;
&lt;br /&gt;
Stevan, this is supposed to be fixed in the transition from the &amp;lt;code&amp;gt;raw&amp;lt;/code&amp;gt; schema to the &amp;lt;code&amp;gt;tidy&amp;lt;/code&amp;gt; schema.  At least this is where the date and time intervals are normalized. &lt;br /&gt;
&lt;br /&gt;
Maybe it got missed, or just needed bringing up as an issue.  No real point in fixing it in MS Access.  Doing so feels like losing information, and possibly being inconsistent in the way time intervals are normalized.&lt;br /&gt;
&lt;br /&gt;
Update 2026-07-22 SRE cannot reproduce this error, e.g., see below querying the tidy schema, which returns zero records; calling this resolved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    &amp;quot;FB_FOL_date&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_begin_feed_time&amp;quot;,&lt;br /&gt;
    &amp;quot;FB_end_feed_time&amp;quot;&lt;br /&gt;
FROM tidy.&amp;quot;FOOD_BOUT&amp;quot;&lt;br /&gt;
WHERE EXTRACT(SECOND FROM &amp;quot;FB_begin_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
   OR EXTRACT(SECOND FROM &amp;quot;FB_end_feed_time&amp;quot;) &amp;lt;&amp;gt; 0&lt;br /&gt;
ORDER BY &amp;quot;FB_FOL_date&amp;quot;, &amp;quot;FB_FOL_B_focal_AnimID&amp;quot;, &amp;quot;FB_begin_feed_time&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#96) There are GROOMINGS rows that have direction values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 26 GROOMINGS rows that have have a direction other than &amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;, which are the only values that map to an allowable direction.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_direction, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_direction),&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_direction;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_direction,&lt;br /&gt;
quote_literal(grm_direction) AS quoted_direction,&lt;br /&gt;
UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_direction, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_direction, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_direction, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;G&amp;#039;, &amp;#039;R&amp;#039;, &amp;#039;M&amp;#039;, &amp;#039;U&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#97) There are GROOMINGS rows that have certainty values that do not match to mappable values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 303 GROOMINGS rows that have have a direction other than &amp;#039;Y&amp;#039; or &amp;#039;N&amp;#039;, which are the only values that map to an allowable certainty (Y -&amp;gt; 1; N -&amp;gt; 0). Only values of 0 and 1 are allowed in events.certainty. &lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;) AS raw_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
GROUP BY&lt;br /&gt;
COALESCE(grm_time_certainty, &amp;#039;&amp;lt;NULL&amp;gt;&amp;#039;),&lt;br /&gt;
quote_literal(grm_time_certainty),&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)),&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;))&lt;br /&gt;
ORDER BY row_count DESC, raw_certainty;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
grm_fol_date,&lt;br /&gt;
grm_fol_b_focal_animid,&lt;br /&gt;
grm_b_partner_animid,&lt;br /&gt;
grm_time_certainty,&lt;br /&gt;
quote_literal(grm_time_certainty) AS quoted_certainty,&lt;br /&gt;
UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS upper_no_trim,&lt;br /&gt;
LENGTH(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
(&lt;br /&gt;
  COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;^[[:space:]]&amp;#039;&lt;br /&gt;
  OR COALESCE(grm_time_certainty, &amp;#039;&amp;#039;) ~ &amp;#039;[[:space:]]$&amp;#039;&lt;br /&gt;
) AS has_edge_space,&lt;br /&gt;
grm_problems,&lt;br /&gt;
grm_comments,&lt;br /&gt;
grm_extracted_by&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE UPPER(COALESCE(grm_time_certainty, &amp;#039;&amp;#039;)) NOT IN (&amp;#039;Y&amp;#039;, &amp;#039;N&amp;#039;)&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== (#98) --- empty placeholder to restore count ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#99) There are GROOMINGS rows for which the extractedBy values is empty. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 10,710 GROOMINGS rows for which the extractedBy values is empty. This value cannot be NULL.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
    grm_fol_date,&lt;br /&gt;
    grm_fol_b_focal_animid,&lt;br /&gt;
    grm_b_partner_animid,&lt;br /&gt;
    grm_direction,&lt;br /&gt;
    grm_time_certainty,&lt;br /&gt;
    grm_extracted_by,&lt;br /&gt;
    quote_literal(grm_extracted_by) AS quoted_extracted_by,&lt;br /&gt;
    LENGTH(COALESCE(grm_extracted_by, &amp;#039;&amp;#039;)) AS raw_length,&lt;br /&gt;
    grm_problems,&lt;br /&gt;
    grm_comments&lt;br /&gt;
FROM clean.groom_bout&lt;br /&gt;
WHERE NULLIF(BTRIM(grm_extracted_by), &amp;#039;&amp;#039;) IS NULL&lt;br /&gt;
ORDER BY grm_fol_date, grm_fol_b_focal_animid, grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#100) There are GROOMINGS rows that have extractedby values that do not match to mappable people values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 14,771 GROOMINGS rows that have extractedby values that do not match to mappable codes.people values.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
==== Summary of non-compliant values ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  COUNT(*) AS row_count&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
GROUP BY BTRIM(gb.grm_extracted_by)&lt;br /&gt;
ORDER BY row_count DESC, normalized_extracted_by;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Full record ====&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT&lt;br /&gt;
  gb.grm_fol_date,&lt;br /&gt;
  gb.grm_fol_b_focal_animid,&lt;br /&gt;
  gb.grm_b_partner_animid,&lt;br /&gt;
  gb.grm_extracted_by,&lt;br /&gt;
  BTRIM(gb.grm_extracted_by) AS normalized_extracted_by,&lt;br /&gt;
  gb.grm_direction,&lt;br /&gt;
  gb.grm_time_certainty,&lt;br /&gt;
  gb.grm_problems,&lt;br /&gt;
  gb.grm_comments&lt;br /&gt;
FROM clean.groom_bout gb&lt;br /&gt;
WHERE NULLIF(BTRIM(gb.grm_extracted_by), &amp;#039;&amp;#039;) IS NOT NULL&lt;br /&gt;
  AND NOT EXISTS (&lt;br /&gt;
    SELECT 1&lt;br /&gt;
    FROM codes.people p&lt;br /&gt;
    WHERE p.person = BTRIM(gb.grm_extracted_by)&lt;br /&gt;
  )&lt;br /&gt;
ORDER BY gb.grm_fol_date, gb.grm_fol_b_focal_animid, gb.grm_b_partner_animid;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== * (#101) There are AGGRESSION_EVENT rows that have NULL focals, a NULL ae_fol_b_focal_id value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  144 aggression_event rows that have NULL for the focal.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_fol_b_focal_id IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown individual, &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt;.&lt;br /&gt;
This means that these rows are converted to WATCHES rows with a Type of &amp;lt;code&amp;gt;B&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;UNK&amp;lt;/code&amp;gt; as the focal.&lt;br /&gt;
&lt;br /&gt;
The data needs to be reviewed to see if this is appropriate.  See also problem #102.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#102) There are AGGRESSION_EVENT rows that have NULL times, a NULL ae_time value ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are  107 aggression_event rows that have NULL for the time.&lt;br /&gt;
Note that all of these rows also have NULL for the focal, problem #101.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT *&lt;br /&gt;
  FROM easy.aggression_event&lt;br /&gt;
  WHERE ae_time IS NULL;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Change these values to the unknown time, midnight, &amp;lt;code&amp;gt;00:00&amp;lt;/code&amp;gt;.&lt;br /&gt;
These need to be reviewed to ensure the change is appropriate.&lt;br /&gt;
Meeting discussions indicate the many, ideally all, of these rows are the result of perusal of a book&amp;#039;s content, where aggressions were mentioned but these aggressions did not otherwise appear in the data.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#103) There are BIOGRAPHY rows that have empty birthgroup values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 64 BIOGRAPHY rows that have empty strings for birthgroup; NULL values are allowed but not empty strings. Some of the offending records are addressed also in Problem #13.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_BirthGroup&amp;quot;) AS birthgroup,&lt;br /&gt;
       length(&amp;quot;B_BirthGroup&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_BirthGroup&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_BGCertainty&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_BirthGroup&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_BirthGroup&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_birthgroup = NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_birthgroup IS DISTINCT FROM NULLIF(BTRIM(b_birthgroup), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#104) There are BIOGRAPHY rows that have empty momid values. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There are 89 BIOGRAPHY rows that have empty strings for momid; NULL values are allowed but not empty strings.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: (temporary) solution applied to clean so query tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT &amp;quot;B_AnimID&amp;quot;,&lt;br /&gt;
       &amp;quot;B_AnimName&amp;quot;,&lt;br /&gt;
       quote_nullable(&amp;quot;B_MomID&amp;quot;) AS momid,&lt;br /&gt;
       length(&amp;quot;B_MomID&amp;quot;) AS character_length,&lt;br /&gt;
       octet_length(&amp;quot;B_MomID&amp;quot;) AS byte_length,&lt;br /&gt;
       &amp;quot;B_Birthdate&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;BIOGRAPHY&amp;quot;&lt;br /&gt;
 WHERE &amp;quot;B_MomID&amp;quot; IS NOT NULL&lt;br /&gt;
       AND BTRIM(&amp;quot;B_MomID&amp;quot;) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY &amp;quot;B_AnimID&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Solution ===&lt;br /&gt;
&lt;br /&gt;
Propose converting empty strings to NULL in the clean schema, PI to confirm.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
UPDATE biography&lt;br /&gt;
  SET b_momid = NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;)&lt;br /&gt;
  WHERE b_momid IS DISTINCT FROM NULLIF(BTRIM(b_momid), &amp;#039;&amp;#039;);&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== * (#105) There is a comm member log record that is missing a description. ==&lt;br /&gt;
&lt;br /&gt;
=== Problem ===&lt;br /&gt;
&lt;br /&gt;
There is a single comm member log records where the description is missing.&lt;br /&gt;
&lt;br /&gt;
=== Bad data ===&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;note: querying tidy&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
SELECT date_of_update,&lt;br /&gt;
       chimp_id,&lt;br /&gt;
       quote_nullable(update_description) AS update_description,&lt;br /&gt;
       length(update_description) AS description_length,&lt;br /&gt;
       update_rationale,&lt;br /&gt;
       &amp;quot;made by&amp;quot;&lt;br /&gt;
  FROM tidy.&amp;quot;COMMUNITY_MEMBERSHIP_UPDATE_LOG&amp;quot;&lt;br /&gt;
 WHERE update_description IS NULL&lt;br /&gt;
       OR BTRIM(update_description) = &amp;#039;&amp;#039;&lt;br /&gt;
 ORDER BY date_of_update,&lt;br /&gt;
          chimp_id,&lt;br /&gt;
          update_rationale,&lt;br /&gt;
          &amp;quot;made by&amp;quot;;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;/div&gt;</summary>
		<author><name>StevanEarl</name></author>
	</entry>
</feed>